Remote Full Time
--
Tahaluf Al Emarat Technical Solutions LLC

Job Details

Job description

Build pipelines that process large volumes of documents (PDF, Word, Excel, and scanned fi les, inEnglish and Arabic) into clean, structured data.

Use LLMs and vision-language models to extract key details into defi ned schemas (structuredoutputs).

Classify documents and rank them against business-defi ned criteria, and explain why each rankingwas given.

Build retrieval over document collections so users and agents can search and ask questions.

Build evaluation:

labelled test sets

precision and recall for extraction and classifi cation

LLM-as-judge where it helps

regression tests whenever prompts or models change

Keep cost and latency under control on self-hosted, open-weight models.

Work with the business users who defi ne the criteria, and turn their feedback into measurableimprovements.

Skills

Must have

6+ years of software engineering, including 1+ year building LLM-powered applications.

Hands-on experience with

document processing

: parsing, OCR, and layout analysis (e.g.Docling, Unstructured, Tesseract, or similar).

Experience building extraction, classifi cation, or ranking systems, and measuring their accuracy.

Strong

Python

.

Working knowledge of the core AI topics below.

Nice to have

Arabic NLP, or experience with multilingual documents.

Vision-language models for document understanding.

Vector search and RAG (e.g. pgvector, Qdrant, OpenSearch).

Fine-tuning small models for extraction or classifi cation.

Learning-to-rank or recommendation systems.

Core AI knowledge

We expect every AI Engineer on our team to be able to discuss these topics with confi dence. For this rolewe expect depth in LLM fundamentals and evaluation, and working knowledge of the rest.

How LLMs work:

the Transformer architecture, tokenisation (and how it affects Arabic text),decoding and sampling, the training pipeline (pre-training, SFT, RLHF / DPO, RL with verifi ablerewards), long-context limits, and common failure modes such as hallucination and promptinjection.

Agentic AI:

agent patterns (ReAct, plan-and-execute, refl ection), tool design, context engineering(retrieval, memory, compaction), and safety controls.

Protocols:

Model Context Protocol (MCP) and Agent2Agent (A2A).

Agent harnesses:

what sits around the model (the loop, tools, permissions, context management),and hands-on experience with at least one agent framework or SDK.

Behavioural vs procedural approaches:

when to let the model decide the steps and when todefi ne them in code, and how to combine the two.

Benchmarks:

what current benchmarks measure and miss, and how to read results critically.

Similar Jobs