School agent Jobs
1385 Jobs Found
<p><h4>Description</h4>
<p><strong>About Maqsam</strong><br>
Maqsam is the leading Arabic AI-powered contact center solution in the MENA region designed to revolutionize customer experience. Leveraging groundbreaking advancements in AI technology, Maqsam has spearheaded accurate Arabic call transcription, advanced analytics, and insights, along with unique AI modeling utilizing its in-house designed LLM.</p>
<p><strong>Your opportunity</strong><br>
As a senior AI agent architect, you will play a key role in designing, optimizing, and scaling AI-powered customer experiences for leading businesses across the region. You will be responsible for translating client requirements, business flows, and knowledge bases into reliable, high-performing AI agents that deliver accurate and natural interactions. This role involves structuring and improving knowledge sources, designing prompts and workflows, evaluating AI agent behavior, and continuously enhancing performance through testing and analysis. As the bridge between technical teams and client needs, you will collaborate closely with product, engineering, and customer success teams to ensure AI solutions are practical, scalable, and aligned with business goals. You will also contribute to defining best practices for AI agent architecture, prompt quality, and conversational design while helping shape the future of Arabic AI experiences at Maqsam.</p>
<h4>Key responsibilities</h4>
<ul>
<li>Review client requirements, knowledge base documents, URLs, and business flows to assess readiness for AI agent implementation.</li>
<li>Structure and optimize knowledge bases to improve retrieval quality, contextual accuracy, and agent reliability.</li>
<li>Design, tune, and iterate on prompts, system messages, and agent instructions to improve response quality.</li>
<li>Set up and improve AI agent workflows, actions, and interaction logic.</li>
<li>Identify gaps, edge cases, and failure points in AI agent behavior and propose practical improvements.</li>
<li>Define testing approaches for AI agents, including stress testing, scenario-based testing, and quality checks.</li>
<li>Review testing results and client feedback, then translate them into clear improvements.</li>
<li>Collaborate with product, engineering, customer success, and other teams to align AI agent behavior with client needs.</li>
<li>Document decisions, issues, and recommendations clearly for internal teams and clients.</li>
<li>Help define best practices for AI agent design, knowledge base structuring, prompt quality, and evaluation.</li>
<li>Support and guide less experienced team members when needed.</li>
</ul>
<h4>Who we’re looking for</h4>
<ul>
<li>Strong technical background with hands-on experience in conversational AI, LLMs, AI agents, or chatbot platforms.</li>
<li>Practical experience in prompt engineering, workflow design, automation systems, or AI agent configuration.</li>
<li>Good understanding of knowledge base structuring, content quality, and how content impacts AI agent performance.</li>
<li>Ability to analyze AI agent behavior, identify root causes of failures, and propose clear solutions.</li>
<li>Experience with testing AI systems, including scenario testing, prompt evaluation, or quality assurance.</li>
<li>Strong critical thinking and attention to detail.</li>
<li>Excellent written and verbal communication skills, especially when explaining technical decisions or giving feedback.</li>
<li>Ability to work effectively with cross-functional teams, including product, engineering, customer success, and client-facing teams.</li>
<li>At least 2–3 years of relevant work experience is preferred.</li>
<li>Experience with RAG systems, MCPs, external tools, APIs, or voice agents is a plus.</li>
<li>Ability to work in a fast-changing environment and take ownership of ambiguous problems.</li>
</ul>
<h4>What we provide you</h4>
<ul>
<li>Join a world-class team at Maqsam, where your growth is our priority.</li>
<li>A collaborative and friendly environment with flexible hours.</li>
<li>Many opportunities to learn, grow, and become a world-class professional.</li>
<li>A fair compensation.</li>
<li>A culture that thrives on open communication and transparency.</li>
<li>First-class health insurance.</li>
<li>Part of something big: work with seasoned professionals and be proud of being part of the exciting world of SaaS business.</li>
</ul></p><p></p>
<p><h4>Description</h4>
<p><strong>About Maqsam</strong><br>
Maqsam is the leading Arabic AI-powered contact center solution in the MENA region designed to revolutionize customer experience. Leveraging groundbreaking advancements in AI technology, Maqsam has spearheaded accurate Arabic call transcription, advanced analytics, and insights, along with unique AI modeling utilizing its in-house designed LLM.</p>
<p><strong>Your opportunity</strong><br>
As a senior AI agent architect, you will play a key role in designing, optimizing, and scaling AI-powered customer experiences for leading businesses across the region. You will be responsible for translating client requirements, business flows, and knowledge bases into reliable, high-performing AI agents that deliver accurate and natural interactions. This role involves structuring and improving knowledge sources, designing prompts and workflows, evaluating AI agent behavior, and continuously enhancing performance through testing and analysis. As the bridge between technical teams and client needs, you will collaborate closely with product, engineering, and customer success teams to ensure AI solutions are practical, scalable, and aligned with business goals. You will also contribute to defining best practices for AI agent architecture, prompt quality, and conversational design while helping shape the future of Arabic AI experiences at Maqsam.</p>
<h4>Key responsibilities</h4>
<ul>
<li>Review client requirements, knowledge base documents, URLs, and business flows to assess readiness for AI agent implementation.</li>
<li>Structure and optimize knowledge bases to improve retrieval quality, contextual accuracy, and agent reliability.</li>
<li>Design, tune, and iterate on prompts, system messages, and agent instructions to improve response quality.</li>
<li>Set up and improve AI agent workflows, actions, and interaction logic.</li>
<li>Identify gaps, edge cases, and failure points in AI agent behavior and propose practical improvements.</li>
<li>Define testing approaches for AI agents, including stress testing, scenario-based testing, and quality checks.</li>
<li>Review testing results and client feedback, then translate them into clear improvements.</li>
<li>Collaborate with product, engineering, customer success, and other teams to align AI agent behavior with client needs.</li>
<li>Document decisions, issues, and recommendations clearly for internal teams and clients.</li>
<li>Help define best practices for AI agent design, knowledge base structuring, prompt quality, and evaluation.</li>
<li>Support and guide less experienced team members when needed.</li>
</ul>
<h4>Who we’re looking for</h4>
<ul>
<li>Strong technical background with hands-on experience in conversational AI, LLMs, AI agents, or chatbot platforms.</li>
<li>Practical experience in prompt engineering, workflow design, automation systems, or AI agent configuration.</li>
<li>Good understanding of knowledge base structuring, content quality, and how content impacts AI agent performance.</li>
<li>Ability to analyze AI agent behavior, identify root causes of failures, and propose clear solutions.</li>
<li>Experience with testing AI systems, including scenario testing, prompt evaluation, or quality assurance.</li>
<li>Strong critical thinking and attention to detail.</li>
<li>Excellent written and verbal communication skills, especially when explaining technical decisions or giving feedback.</li>
<li>Ability to work effectively with cross-functional teams, including product, engineering, customer success, and client-facing teams.</li>
<li>At least 2–3 years of relevant work experience is preferred.</li>
<li>Experience with RAG systems, MCPs, external tools, APIs, or voice agents is a plus.</li>
<li>Ability to work in a fast-changing environment and take ownership of ambiguous problems.</li>
</ul>
<h4>What we provide you</h4>
<ul>
<li>Join a world-class team at Maqsam, where your growth is our priority.</li>
<li>A collaborative and friendly environment with flexible hours.</li>
<li>Many opportunities to learn, grow, and become a world-class professional.</li>
<li>A fair compensation.</li>
<li>A culture that thrives on open communication and transparency.</li>
<li>First-class health insurance.</li>
<li>Part of something big: work with seasoned professionals and be proud of being part of the exciting world of SaaS business.</li>
</ul></p><p></p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<p>We’re Hiring: Real Estate Agent</p><p><br></p><p>We’re looking for a motivated and proactive Real Estate Agent to join our team.</p><p><br></p><p>Requirements:</p><p><br></p><p>* Fluent in English (spoken and written)</p><p>* Minimum 3 years of experience in real estate</p><p>* Dynamic, proactive, and results-driven</p><p>* Strong communication and negotiation skills</p><p><br></p><p>Key Responsibilities:</p><p><br></p><p>* Create and manage new property listings</p><p>* Contact companies, embassies, NGOs, and other organizations to generate tenant leads</p><p>* Build relationships with landlords and potential clients</p><p>* Follow up with inquiries and close rental deals</p><p><br></p><p>If you believe you’re the right fit, we’d love to hear from you. </p> </div><h2 class="h5">Skills</h2>
<div data-jb-field="skills"><p>* Fluent in English (spoken and written)</p><p>* Minimum 3 years of experience in real estate</p><p>* Dynamic, proactive, and results-driven</p><p>* Strong communication and negotiation skills</p></div>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>POSITION SUMMARY Process all reservation requests, changes, and cancellations received by phone, fax, or mail. Identify guest reservation needs and determine appropriate room type. Verify availability of room type and rate. Explain guarantee, special rate, and cancellation policies to callers. Accommodate and document special requests. Answer questions about property facilities/services and room accommodations. Follow sales techniques to maximize revenue. Input and access data in reservation system. Indicate special room reservation types (e.g., complimentary rooms, employee discounts, travel agent inspection rates, and wholesale reservations) by inputting the correct code and rate into the reservation system. Follow proper escalation procedures when addressing guest concerns. Follow all company policies and procedures; ensure uniform and personal appearance are clean and professional; maintain confidentiality of proprietary information; protect company assets; protect the privacy and security of guests and coworkers. Welcome and acknowledge all guests according to company standards; anticipate and address guests service needs; assist individuals with disabilities; thank guests with genuine appreciation. Speak with others using clear and professional language; answer telephones using appropriate etiquette. Develop and maintain positive working relationships with others; support team to reach common goals; listen and respond appropriately to the concerns of other employees. Comply with quality assurance expectations and standards. Move, lift, carry, push, pull, and place objects weighing less than or equal to 10 pounds without assistance. Perform other reasonable job duties as requested by Supervisors.</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><ul><li>Education: High school diploma or G.E.D. equivalent.</li><li>Related Work Experience: No related work experience.</li><li>Supervisory Experience: No supervisory experience.</li><li>License or Certification: None</li></ul><p></p></section>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Position summary</h4>
<p>Process all reservation requests, changes, and cancellations received by phone, fax, or mail. Identify guest reservation needs and determine appropriate room type. Verify availability of room type and rate. Explain guarantee, special rate, and cancellation policies to callers. Accommodate and document special requests. Answer questions about property facilities/services and room accommodations. Follow sales techniques to maximize revenue. Input and access data in reservation system. Indicate special room reservation types (e.g., complimentary rooms, employee discounts, travel agent inspection rates, and wholesale reservations) by inputting the correct code and rate into the reservation system. Follow proper escalation procedures when addressing guest concerns.</p>
<p>Follow all company policies and procedures; ensure uniform and personal appearance are clean and professional; maintain confidentiality of proprietary information; protect company assets; protect the privacy and security of guests and coworkers. Welcome and acknowledge all guests according to company standards; anticipate and address guests’ service needs; assist individuals with disabilities; thank guests with genuine appreciation</p><p></p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<strong><u>JOB PURPOSE</u></strong><br>The Imports/Export agent is responsible on the daily import or export activities at AMM GATEWAY operations, reporting to different shifts, AMM GATEWAY agent will be holding tasks related to inbound flights recovery/clearance or outbound flights export. The agent is also responsible in helping to achieve Gateway KPI’s through effectively working together as a team with the rest of the Gateway staffs and while holding different tasks depending on shifts roster. <br><br><strong><u>PRINCIPAL ACCOUNTABILITIES</u></strong><br><ul><li>Provide a high level of customer service while processing inbound or outbound tasks, taking into account the consideration that all DHL customers have an express requirement and are looking for instant and immediate action.</li><li>Follow Gateway inbound clearance procedures as outlined in the GATEWAY manual to comply with the GSOP procedures and safe working practices.</li><li>Maintain a clean desk policy and file all documents required for storing paper works in a systematic way which are easy to access when required.</li><li>Daily updates on clearance status of all shipments held in AMM GATEWAY for clearance using different DHL applications.</li><li>Daily check points to be created for network visibility using appropriate exception codes for different activities at AMM GATEWAY.</li><li>Maintain a thorough knowledge of all departments, DHL network, products and services so that customers are provided accurate information on transit time, clearance delays, custom paperwork requirements, packing, accounting and sales queries with confidence at all times.</li><li>Highlight any recurring problems that are manifested through traces and then direct the information accordingly so that corrective actions can be taken promptly.</li><li>Work effectively both individually and as part of a team to achieve both individual and department goals and objectives and strive consistently to promote a positive team spirit by learning all activities within the department.</li></ul><br><strong><u>NATURE AND SCOPE</u></strong><br><strong>Context</strong><br>The Gateway import agent department has been structured to provide telephone assistance to customers (internal and external) when making inquiries, logging in customer requests, assisting customers on clearance requirements and inquiries with regards to delayed or missing shipments. The department expands DHL’s capabilities, both re-actively and pro-actively, to ensure that all customers (internal and external) experience consistently high levels of service at all times. <br>The multifunctional Gateway import Agent must ensure that service levels are constantly met and exceeded through accurate updates, trace escalations and regular communication with both the network and the customer. <br><strong>Reporting Relationships</strong><br>Reports directly to the Gateway supervisor.<br><strong>Contact</strong><br>External<br><ul><li>Customers</li><li>DHL Network</li></ul><br>Internal <br><ul><li>Ground Operation</li><li>Customer Care </li><li>CS Manager & Supervisors</li><li>Air operations </li><li>Sales</li><li>Accounts</li><li>Customer Accounting</li><li>Service center T/L</li></ul><br><strong>Problem Solving</strong><br>The job holder must take ownership of all customer inquiries and queries and provide alternatives and solutions closing it out at their end. Traces must be logged and followed up regularly to deliver optimum levels of service at all time. For escalations the jobholder must follow the network standards and consult with the Team Leader/Supervisor if unsure.<br><strong>Decision Making</strong><br>The jobholder will be responsible for giving accurate clearance information to customs agents as instructed by the importer with the ultimate objective being service excellence and customer satisfaction. They must communicate all delays and other service issues to the Team leader/ Supervisor.<br><strong>Planning and Organization</strong><br><br>The job holder must be highly organized in while holding tasks at AMM GATEWAY, daily follow – ups and plan a course of action to ensure that set targets and goals are achieved consistently.<br><br><strong>Job Challenge</strong><br><br>The jobholder must always be pro-active and energetic when dealing with customers efficiency in work practices and effective and friendly communication will be key to delighting our customers, adopt DHL “Right from first time” principle.<br>The jobholder must constantly strive to keep self-updated on all DHL processes, systems, products, technology and set and strive to achieve higher standards, ready to work on different shifts and Gateway operations task.<br><strong>KNOWLEDGE, SKILLS AND EXPERIENCE</strong><br><ul><li>Sound educational back ground with knowledge of the Service Industry, an added advantage.</li><li>Sound knowledge of custom clearance procedures. </li><li>Working knowledge of Microsoft Word, Excel and Power Point.</li><li>Good oral and written communication skills – English & Arabic. </li><li>Self motivated individual capable of taking ownership and working independently.</li><li>Tolerance for stress in a fast paced working environment.</li><li>Excellent planning and organising skills.</li><li>Passion for delighting customers.</li><li>Good team player.</li><li>Adheres to policies and procedures.</li><li>Possesses good relationship building and interpersonal skills.</li></ul><br><br> </div>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>POSITION SUMMARY Process all reservation requests, changes, and cancellations received by phone, fax, or mail. Identify guest reservation needs and determine appropriate room type. Verify availability of room type and rate. Explain guarantee, special rate, and cancellation policies to callers. Accommodate and document special requests. Answer questions about property facilities/services and room accommodations. Follow sales techniques to maximize revenue. Input and access data in reservation system. Indicate special room reservation types (e.g., complimentary rooms, employee discounts, travel agent inspection rates, and wholesale reservations) by inputting the correct code and rate into the reservation system. Follow proper escalation procedures when addressing guest concerns. Follow all company policies and procedures; ensure uniform and personal appearance are clean and professional; maintain confidentiality of proprietary information; protect company assets; protect the privacy and security of guests and coworkers. Welcome and acknowledge all guests according to company standards; anticipate and address guests service needs; assist individuals with disabilities; thank guests with genuine appreciation. Speak with others using clear and professional language; answer telephones using appropriate etiquette. Develop and maintain positive working relationships with others; support team to reach common goals; listen and respond appropriately to the concerns of other employees. Comply with quality assurance expectations and standards. Move, lift, carry, push, pull, and place objects weighing less than or equal to 10 pounds without assistance. Perform other reasonable job duties as requested by Supervisors.</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><p>Education: High school diploma or G.E.D. equivalent.</p><p>Related Work Experience: No related work experience.</p><p>Supervisory Experience: No supervisory experience.</p><p>License or Certification: None</p><p></p></section>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>POSITION SUMMARY Process all reservation requests, changes, and cancellations received by phone, fax, or mail. Identify guest reservation needs and determine appropriate room type. Verify availability of room type and rate. Explain guarantee, special rate, and cancellation policies to callers. Accommodate and document special requests. Answer questions about property facilities/services and room accommodations. Follow sales techniques to maximize revenue. Input and access data in reservation system. Indicate special room reservation types (e.g., complimentary rooms, employee discounts, travel agent inspection rates, and wholesale reservations) by inputting the correct code and rate into the reservation system. Follow proper escalation procedures when addressing guest concerns. Follow all company policies and procedures; ensure uniform and personal appearance are clean and professional; maintain confidentiality of proprietary information; protect company assets; protect the privacy and security of guests and coworkers. Welcome and acknowledge all guests according to company standards; anticipate and address guests service needs; assist individuals with disabilities; thank guests with genuine appreciation. Speak with others using clear and professional language; answer telephones using appropriate etiquette. Develop and maintain positive working relationships with others; support team to reach common goals; listen and respond appropriately to the concerns of other employees. Comply with quality assurance expectations and standards. Move, lift, carry, push, pull, and place objects weighing less than or equal to 10 pounds without assistance. Perform other reasonable job duties as requested by Supervisors.</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><p>Education: high school diploma or G.E.D. equivalent.</p><p>Related Work Experience: No related work experience.</p><p>Supervisory Experience: No supervisory experience.</p><p>License or Certification: None</p><p></p></section>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply? Pass qualification(s)? Join a project? Complete tasks? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply \u2192 Pass qualification(s) \u2192 Join a project \u2192 Complete tasks \u2192 Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply \u2192 Pass qualification(s) \u2192 Join a project \u2192 Complete tasks \u2192 Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>