Description
About the opportunity
We are hiring on behalf of a major government transformation initiative in Abu Dhabi that is building one of the world’s most ambitious applied AI programs. This organisation is developing the infrastructure and platforms that enable AI-powered public services at scale. The work spans sovereign model serving, vector infrastructure, retrieval systems, developer tooling, and deployment platforms that support critical government applications. This is a rare opportunity to build foundational AI infrastructure with significant real-world impact, in an environment where performance, reliability, and security are mission-critical.
Role overview
We are seeking a Staff AI Platform Engineer to own the platform infrastructure that enables engineering teams to build and operate AI systems efficiently and reliably. This is a senior individual contributor role focused on model serving, vector infrastructure, data pipelines, observability, and deployment platforms. Your success will be measured by the productivity and reliability gains you enable for other engineers.
You will combine deep platform engineering expertise with a strong understanding of AI workload requirements, including inference performance, retrieval systems, and operational complexity.
At the staff level, you will:
- Define and evolve the technical foundations of the AI platform
- Build systems that accelerate engineering teams across the organisation
- Set standards for reliability, security, and scalability
- Lead architecture decisions and evaluate emerging technologies
- Use AI coding tools such as Codex, Claude Code, or similar as part of your daily workflow
Key responsibilities
AI platform infrastructure
Design and operate model serving infrastructure using vLLM, TGI, TensorRT-LLM, or similar technologies.
Build and manage vector infrastructure, embedding pipelines, and retrieval systems.
Develop data pipelines for document ingestion, transformation, and storage.
Own deployment platforms, CI/CD pipelines, and infrastructure-as-code.
Observability & reliability
Build observability across the AI stack, including latency, throughput, and model behavior monitoring.
Define SLOs, perform capacity planning, and lead reliability engineering initiatives.
Lead incident response, root-cause analysis, and postmortems.
Establish platform standards for operational excellence.
Force multiplication
Build reusable internal tooling, SDKs, and deployment patterns.
Partner with engineering teams to translate infrastructure needs into scalable platform capabilities.
Mentor engineers and raise technical standards across the organisation.
Basic qualifications
- 10+ years of platform, infrastructure, or backend engineering experience
- Proven experience operating at staff or principal engineer level
- Deep expertise in Azure (AWS or GCP also valued)
- Strong Kubernetes and Docker experience
- Hands-on experience with Terraform, CI/CD pipelines, and infrastructure-as-code
- Experience building production data pipelines
- Proficiency in Python, Java/Kotlin, or Go
- Strong PostgreSQL knowledge at production scale
- Experience with distributed tracing, metrics, logging, and alerting
- Practical use of AI coding assistants such as Codex or Claude Code
- Excellent written and verbal communication skills
Preferred qualifications
- Experience building RAG systems and retrieval infrastructure
- Experience with LangGraph, LangChain, Semantic Kernel, or similar frameworks
- Hands-on experience with GPU-based inference infrastructure
- Experience with vector databases such as pgvector, Qdrant, Pinecone, or Weaviate
- Familiarity with LLM inference optimization and evaluation frameworks
- Experience with speech and conversational AI systems
- Experience building internal developer platforms
- Security and compliance experience in highly regulated environments
Technology stack
Backend & platform: Python, Java, Kotlin, Go, REST, gRPC
AI infrastructure: vLLM, TGI, TensorRT-LLM, GPU Serving
AI & LLM: LangChain, LangGraph, Microsoft Agent Framework, vector databases
Data: PostgreSQL, Redis, DocumentDB, Azure Blob Storage
Infrastructure: Azure, Docker, Kubernetes, Terraform, CI/CD
Observability: Langfuse, Grafana, Prometheus, OpenTelemetry
Security: Entra ID, RBAC, Secrets Management, Zero Trust
AI development: Codex, Claude Code, OpenCode
What we’re looking for
You are an engineer who:
- Thinks in systems and enables others to move faster
- Designs infrastructure that is reliable, scalable, and easy to operate
- Grounds technical decisions in evidence and operational data
- Raises engineering standards through architecture and mentorship
- Communicates clearly and transparently
- Embraces AI-native development practices
Why apply?
This role offers the opportunity to shape the technical foundations of a world-class AI platform supporting transformative public services in Abu Dhabi. You will tackle complex infrastructure challenges at the intersection of AI, cloud computing, and platform engineering while delivering meaningful impact at a national scale.
الوصف
عن الفرصة
نحن نوظف نيابة عن مبادرة تحول حكومي كبرى في أبوظبي تبني أحد أكثر برامج الذكاء الاصطناعي التطبيقي طموحاً في العالم. تقوم هذه المنظمة بتطوير البنية التحتية والمنصات التي تمكّن الخدمات العامة المدعومة بالذكاء الاصطناعي على نطاق واسع. تمتد الأعمال عبر تشغيل النماذج السيادية، والبنية التحتية للنواقل، وأنظمة الاسترجاع، وأدوات المطورين، ومنصات النشر التي تدعم التطبيقات الحكومية الحرجة. هذه فرصة نادرة لبناء بنية أساسية للذكاء الاصطناعي لها أثر حقيقي في العالم، في بيئة تكون فيها الأداء والموثوقية والأمن أموراً حيوية للمهمة.
نظرة على الدور
نبحث عن مهندس منصة ذكاء اصطناعي وظيفته امتلاك بنية المنصة التي تمكّن فرق الهندسة من بناء وتشغيل أنظمة ذكاء اصطناعي بكفاءة وموثوقية. هذا دور مساهم فردي كبير يتركز على تشغيل النماذج، وبنية تحتية للنواقل، وأ pipelines البيانات، والرصد، ومنصات النشر. سيتم قياس نجاحك من خلال زيادة الإنتاجية والموثوقية التي تتيحها للمهندسين الآخرين.
ستجمع بين خبرة عميقة في هندسة المنصة وفهم قوي لمتطلبات أعباء الذكاء الاصطناعي، بما في ذلك أداء الاستدلال وأنظمة الاسترجاع والتعقيد التشغيلي.
على مستوى الموظف، ستقوم بـ:
- تحديد وتطوير الأسس التقنية لمنصة الذكاء الاصطناعي
- بناء أنظمة تسرّع فرق الهندسة عبر المؤسسة
- وضع معايير للموثوقية والأمن وقابلية التوسع
- قيادة قرارات الهندسة المعمارية وتقييم التقنيات الناشئة
- استخدام أدوات برمجة الذكاء الاصطناعي مثل Codex أو Claude Code أو ما يشابهه كجزء من سير عملك اليومي
المهام الأساسية
بنـية منصة الذكاء الاصطناعي
صمم وشغّل بنية تشغيل النماذج باستخدام vLLM، TGI، TensorRT-LLM، أو تقنيات مماثلة.
ابنِ وأدر بنية البيانات المتجهة، وخطوط إدراج المصطلحات، وأنظمة الاسترجاع.
طور خطوط بيانات لاستيعاب المستندات والتحويل والتخزين.
امتلك منصات النشر، وخطوط CI/CD، وبنية التحتية ككود.
الرصد والموثوقية
ابنِ الرصد عبر تكدس الذكاء الاصطناعي، بما في ذلك الكمون، والإنتاجية، ومراقبة سلوك النموذج.
ضع SLOs، وأجر تخطيط السعة، وقُد مبادرات هندسة الموثوقية.
قُد استجابة الحوادث، وتحليل السبب الجذري، وما بعد الحادثة.
أسس معايير المنصة للتميز التشغيلي.
ضربات القوة
ابن أدوات داخلية قابلة لإعادة الاستخدام، وحزم تطوير البرمجيات، ونماذج النشر.
شارك مع فرق الهندسة لتحويل احتياجات البنية التحتية إلى قدرات منصة قابلة للتوسع.
قم بتوجيه المهندسين ورفع المعايير التقنية عبر المؤسسة.
المؤهلات الأساسية
- 10+ سنوات من الخبرة في المنصة أو البنية التحتية أو هندسة الخلفية
- خبرة مثبتة في العمل كمستوى موظف أو مهندس رئيسي
- خبرة عميقة في Azure (AWS أو GCP قيم أيضاً)
- خبرة قوية في Kubernetes و Docker
- خبرة عملية مع Terraform، وخطوط CI/CD، والبنية التحتية ككود
- خبرة في بناء خطوط بيانات إنتاجية
- إتقان Python أو Java/Kotlin أو Go
- معرفة قوية بـ PostgreSQL على نطاق الإنتاج
- خبرة في التتبع الموزع، والقياسات، والسجلات، والتنبيه
- استخدام عملي لمساعدي ترميز الذكاء الاصطناعي مثل Codex أو Claude Code
- مهارات اتصال مكتوبة وشفوية ممتازة
المؤهلات المفضلة
- خبرة في بناء أنظمة RAG وبنية الاسترجاع
- خبرة مع LangGraph، LangChain، Semantic Kernel، أو أطر مشابهة
- خبرة عملية في بنية استدلال قائمة على GPU
- خبرة مع قواعد بيانات المتجهات مثل pgvector، Qdrant، Pinecone، أو Weaviate
- familiar with LLM inference optimization and evaluation frameworks
- خبرة في أنظمة الكلام والذكاء الاصطناعي المحادثي
- خبرة في بناء منصات مطورين داخلية
- خبرة الأمن والامتثال في بيئات عالية التنظيم
تكدس التقنية
الخلفية والخلفية المنصة: Python, Java, Kotlin, Go, REST, gRPC
بنية الذكاء الاصطناعي: vLLM, TGI, TensorRT-LLM, GPU Serving
الذكاء الاصطناعي وLLM: LangChain, LangGraph, Microsoft Agent Framework, قواعد بيانات المتجهات
البيانات: PostgreSQL, Redis, DocumentDB, Azure Blob Storage
البنية التحتية: Azure, Docker, Kubernetes, Terraform, CI/CD
الرصد: Langfuse, Grafana, Prometheus, OpenTelemetry
الأمن: Entra ID, RBAC, Secrets Management, Zero Trust
تطوير الذكاء الاصطناعي: Codex, Claude Code, OpenCode
ما الذي نبحث عنه
أنت مهندس يفعل ما يلي:
- يفكر في الأنظمة ويمكن الآخرين من التحرك بسرعة أكبر
- يصمم بنية تحتية موثوقة وقابلة للتوسع وسهلة التشغيل
- يؤسس القرارات التقنية بالأدلة وبيانات التشغيل
- يرفع معايير الهندسة من خلال الهندسة المعمارية والتوجيه
- يتواصل بوضوح وشفافية
- يتبنّى ممارسات التطوير المعتمدة على الذكاء الاصطناعي
لماذا التقديم؟
تتيح هذه المرحلة فرصة تشكيل الأسس التقنية لمنصة ذكاء اصطناعي عالمية المستوى تدعم خدمات عامة تحويلية في أبوظبي. ستتعامل مع تحديات بنية تحتية معقدة عند تقاطع الذكاء الاصطناعي، والحوسبة السحابية، وهندسة المنصات مع تقديم أثر حقيقي على المستوى الوطني.