نظرة عامة على الدور
نحن نبحث عن مهندس بنية تحتية ومنصة ذكاء اصطناعي أول/راقي المهارة للانضمام إلى فريق عميلنا في الرياض. في هذا الدور، ستكون مسؤولاً عن بناء وإدارة وتحسين بنية تحتية للذكاء الاصطناعي قابلة للتوسع وبيئات الحوسبة التي تدعم أحمال العمل عالية الأداء، بما في ذلك خطوط أنابيب AI/ML المعززة بالـ GPU، جدولة العناقيد، والتنسيق.
المسؤوليات الأساسية
- نشر، صيانة، وتحسين عنقودات الحوسبة المعتمدة على GPU والبنية التحتية.
- إدارة وتشغيل أدوات ومنصات تنسيق الـ GPU مثل:
- مدير الأوامر الأساسي من Nvidia (ضروري)
- حزمة Nvidia AI Enterprise
- مشغّلات GPU/شبكة Nvidia
- NIMs وBlueprints من Nvidia
- تكوين، نشر، وصيانة أحمال عمل الحوسبة باستخدام أدوات الجدولة والتنسيق بما في ذلك:
- Slurm (ضروري)
- Kubernetes عادي
- تثبيت، تكوين، وصيانة نظام التشغيل الأساسي (مثلاً Canonical Ubuntu) وبرمجيات النظام الداعمة.
- مراقبة وتتبع أداء البنية التحتية والتوافر والموثوقية؛ ضمان وقت تشغيل عالي لحِزَم AI/ML.
- التعاون مع علماء البيانات ومهندسي ML وفرق التطوير لتحديد متطلبات البنية التحتية وتخصيص الموارد وتدفقات النشر.
- تطوير نصوص أتمتة وعمليات CI/CD وأفضل الممارسات لتوفير البنية التحتية وإدارتها.
- توثيق الهندسة والتكوينات والإجراءات التشغيلية؛ فرض سياسات الأمان والامتثال والنسخ الاحتياطي.
المتطلبات
المهارات والخبرة المطلوبة
- خبرة مثبتة في إدارة بنية AI/ML قائمة على GPU وعنقودات الحوسبة.
- خبرة عملية مع:
- مدير الأوامر الأساسي من Nvidia
- Nvidia AI Enterprise Suite
- مشغّلات GPU/شبكة Nvidia، NIMs، Blueprints
- خبرة قوية مع التنسيق عبر Slurm و/أو Kubernetes.
- مهارات إدارة نظام Linux صلبة — ويفضل على Ubuntu أو توزيعات مشابهة.
- قدرة قوية على التهيئة/الأتمتة (مثلاً Bash أو Python أو أدوات ذات صلة) لتوفير، نشر، وصيانة.
- مهارات ممتازة في استكشاف الأخطاء وتحسين الأداء.
- الخبرة في التعاون مع فرق ML/علم البيانات ودمج البنية التحتية مع سير عملهم.
- فهم قوي لشبكات الحاسوب والأمان وتخصيص الموارد وأفضل الممارسات لإدارة العنقود.
المؤهلات المفضلة
- خبرة سابقة في العمل ضمن فريق بنية تحتية عالية الأداء (HPC) أو موجهة للذكاء الاصطناعي.
- معرفة بالحاويات، وتنظيم الحاويات، و GPUs في بيئات سحابية أو محلية.
- خبرة في CI/CD، والبنية التحتية كرمز (مثل Terraform، Ansible)، وأدوات المراقبة وتسجيل السجلات.
- الإلمام بجدولة الأحمال، وطوابير الوظائف، وحدود الموارد، وبيئات مشاركة الـ GPU.
Role overview
We are seeking a highly skilled senior AI infrastructure and platform engineer to join our client’s team in Riyadh. In this role, you’ll be responsible for building, managing, and optimizing scalable AI infrastructure and compute environments that support high-performance workloads, including GPU-accelerated AI/ML pipelines, cluster scheduling, and orchestration.
Key responsibilities
- Deploy, maintain, and optimize GPU-based compute clusters and infrastructure.
- Manage and operate GPU orchestration tools and platforms such as:
- Nvidia Base Command Manager (critical)
- Nvidia AI Enterprise Suite
- Nvidia GPU and Network Operators
- Nvidia NIMs and Blueprints
- Configure, deploy, and maintain compute workloads using scheduling and orchestration tools including:
- Slurm (critical)
- Vanilla Kubernetes
- Install, configure, and maintain the underlying OS (e.g. Canonical Ubuntu) and supporting system software.
- Monitor and troubleshoot infrastructure performance, availability, and reliability; ensure high uptime for AI/ML workloads.
- Work with data scientists, ML engineers, and development teams to define infrastructure requirements, resource allocation, and deployment workflows.
- Develop automation scripts, CI/CD pipelines, and best practices for infrastructure provisioning and management.
- Document architecture, configurations, and operational procedures; enforce security, compliance, and backup policies.
Requirements
Required skills & experience
- Proven experience managing GPU-based AI/ML infrastructure and compute clusters.
- Hands-on experience with:
- Nvidia Base Command Manager
- Nvidia AI Enterprise Suite
- Nvidia GPU/Network Operators, NIMs, Blueprints
- Strong experience with Slurm and/or Kubernetes orchestration.
- Solid Linux system administration skills — preferably on Ubuntu or similar distributions.
- Strong scripting/automation ability (e.g. Bash, Python, or relevant tooling) for provisioning, deployment, and maintenance.
- Excellent troubleshooting and performance-tuning skills.
- Experience collaborating with ML/data science teams and integrating infrastructure with their workflows.
- Strong understanding of networking, security, resource allocation, and cluster management best practices.
Preferred qualifications
- Previous experience working in a high-performance computing (HPC) or AI-focused infrastructure team.
- Knowledge of containerization, container orchestration, and GPUs in cloud or on-prem environments.
- Experience with CI/CD, infrastructure-as-code (e.g. Terraform, Ansible), monitoring tools, and logging setups.
- Familiarity with workload scheduling, job queuing, resource quotas, and GPU-shared environments.