Real Estate Sales Jobs in Jordan
1646 Jobs Found
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> Requirements and benefits Educational qualifications A Master’s Degree in Computer Science, Software Engineering, Data Science / Data Analytics, Artificial Intelligence / Machine Learning, Computational Linguistics / Natural Language Processing (NLP), Information Systems or other related fields.<br> Bachelor’s degree is accepted if only candidate has 5 years of experience in the field.<br> Academic and/or Professional Experience Candidates should have a minimum of 3 years of professional experience in related roles or domain - specifically for QA-automation/testing or cybersecurity roles How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Compensation: Paid per accepted task.<br> Your rate depends on the qualification tier you reach and how efficiently you complete tasks — up to the equivalent of $50/hr .<br> Because payment is per task, a faster pace raises your effective hourly rate.<br> Why this freelance opportunity might be a great fit for you?<br> Take part in a part-time, remote, freelance project that fits around your primary professional or academic commitments.<br> Work on advanced AI projects and gain valuable experience that enhances your portfolio.<br> - Influence how future AI models understand and communicate in your field of expertise.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<p><h4>Description</h4>
<p>Robusta Technology Group (RTG) is a key driver of digital transformation by providing a holistic tech ecosystem. RTG works with its local and international partners to help build digital customer experiences, establish engineering hubs, and build ventures across multiple industries and domains. In this pursuit, RTG serves as a catalyst for impact and growth through events, spaces, and content focused on creating impact and growth across the different interactions.</p>
<p>Octopus is proud to be part of the Robusta Technology Group (RTG), a leading tech consultancy group. With a decade of experience and a successful track record of delivering over 300 projects across Europe, the Middle East, and North America, RTG has established itself as a preferred employer in the Egyptian market. Octopus and Robusta are building a bridge between Europe and Africa, creating tailored hub solutions to connect companies with top talent across the globe.</p>
<h4>Job overview</h4>
<p>We are seeking a highly skilled and experienced Salesforce Marketing Cloud Senior (Next) professional to lead the design, implementation, optimization, and management of customer engagement and marketing automation solutions using Salesforce Marketing Cloud and Marketing Cloud Next capabilities. The ideal candidate will have deep expertise in omnichannel campaign execution, customer journey orchestration, data-driven marketing strategies, and advanced Salesforce ecosystem integrations. This role requires both strategic thinking and hands-on technical execution to deliver scalable and personalized customer experiences.</p>
<h4>Role responsibilities</h4>
<p><strong>Platform management & solution design</strong><br>
Lead the implementation and optimization of Salesforce Marketing Cloud solutions, including Marketing Cloud Next functionalities. Design scalable marketing automation architectures aligned with business and customer engagement goals. Configure and manage core SFMC modules such as:</p>
<ul>
<li>Journey Builder</li>
<li>Email Studio</li>
<li>Automation Studio</li>
<li>Mobile Studio</li>
<li>Advertising Studio</li>
<li>Data Cloud integrations</li>
<li>Personalization and AI-driven capabilities</li>
</ul>
<p><strong>Campaign management</strong><br>
Develop, execute, and optimize multi-channel marketing campaigns across email, SMS, push notifications, and digital channels. Build and maintain customer journeys with advanced segmentation and personalization strategies. Ensure campaign delivery accuracy, quality assurance, and performance tracking.</p>
<p><strong>Data & integration</strong><br>
Integrate Salesforce Marketing Cloud with CRM, CDP, APIs, and third-party platforms. Manage data extensions, SQL queries, automations, and audience segmentation. Support real-time data synchronization and customer lifecycle management.</p>
<p><strong>Technical development</strong><br>
Develop dynamic and personalized content using:</p>
<ul>
<li>AMPscript</li>
<li>SSJS</li>
<li>HTML/CSS</li>
<li>SQL</li>
</ul>
<p>Build reusable templates, automation workflows, and scalable campaign frameworks. Troubleshoot platform issues and optimize system performance.</p>
<p><strong>Analytics & optimization</strong><br>
Monitor campaign KPIs and customer engagement metrics. Generate insights and recommendations to improve conversion, retention, and customer experience. Support A/B testing, reporting dashboards, and marketing performance analysis.</p>
<p><strong>Leadership & collaboration</strong><br>
Collaborate with marketing, CRM, sales, analytics, and technology teams. Mentor junior team members and provide SFMC best practice guidance. Participate in solution architecture discussions and stakeholder workshops.</p>
<h4>Requirements</h4>
<ul>
<li>Bachelor’s degree in Computer Science, Information Systems, Marketing, or related field.</li>
<li>5+ years of hands-on experience with Salesforce Marketing Cloud.</li>
<li>Strong experience with Marketing Cloud Next capabilities and modern customer engagement platforms.</li>
<li>Proven expertise in:</li>
<ul>
<li>Journey Builder</li>
<li>Automation Studio</li>
<li>Email Studio</li>
<li>SQL</li>
<li>AMPscript</li>
<li>API integrations</li>
</ul>
<li>Experience integrating SFMC with Salesforce CRM, Data Cloud, and CDP solutions.</li>
<li>Strong understanding of customer lifecycle marketing and omnichannel strategies.</li>
<li>Excellent analytical, troubleshooting, and communication skills.</li>
</ul></p><p></p>
<p><h4>Description</h4>
<p>Robusta Technology Group (RTG) is a key driver of digital transformation by providing a holistic tech ecosystem. RTG works with its local and international partners to help build digital customer experiences, establish engineering hubs, and build ventures across multiple industries and domains. In this pursuit, RTG serves as a catalyst for impact and growth through events, spaces, and content focused on creating impact and growth across the different interactions.</p>
<p>Octopus is proud to be part of the Robusta Technology Group (RTG), a leading tech consultancy group. With a decade of experience and a successful track record of delivering over 300 projects across Europe, the Middle East, and North America, RTG has established itself as a preferred employer in the Egyptian market. Octopus and Robusta are building a bridge between Europe and Africa, creating tailored hub solutions to connect companies with top talent across the globe.</p>
<h4>Job overview</h4>
<p>We are seeking a highly skilled and experienced Salesforce Marketing Cloud Senior (Next) professional to lead the design, implementation, optimization, and management of customer engagement and marketing automation solutions using Salesforce Marketing Cloud and Marketing Cloud Next capabilities. The ideal candidate will have deep expertise in omnichannel campaign execution, customer journey orchestration, data-driven marketing strategies, and advanced Salesforce ecosystem integrations. This role requires both strategic thinking and hands-on technical execution to deliver scalable and personalized customer experiences.</p>
<h4>Role responsibilities</h4>
<p><strong>Platform management & solution design</strong><br>
Lead the implementation and optimization of Salesforce Marketing Cloud solutions, including Marketing Cloud Next functionalities. Design scalable marketing automation architectures aligned with business and customer engagement goals. Configure and manage core SFMC modules such as:</p>
<ul>
<li>Journey Builder</li>
<li>Email Studio</li>
<li>Automation Studio</li>
<li>Mobile Studio</li>
<li>Advertising Studio</li>
<li>Data Cloud integrations</li>
<li>Personalization and AI-driven capabilities</li>
</ul>
<p><strong>Campaign management</strong><br>
Develop, execute, and optimize multi-channel marketing campaigns across email, SMS, push notifications, and digital channels. Build and maintain customer journeys with advanced segmentation and personalization strategies. Ensure campaign delivery accuracy, quality assurance, and performance tracking.</p>
<p><strong>Data & integration</strong><br>
Integrate Salesforce Marketing Cloud with CRM, CDP, APIs, and third-party platforms. Manage data extensions, SQL queries, automations, and audience segmentation. Support real-time data synchronization and customer lifecycle management.</p>
<p><strong>Technical development</strong><br>
Develop dynamic and personalized content using:</p>
<ul>
<li>AMPscript</li>
<li>SSJS</li>
<li>HTML/CSS</li>
<li>SQL</li>
</ul>
<p>Build reusable templates, automation workflows, and scalable campaign frameworks. Troubleshoot platform issues and optimize system performance.</p>
<p><strong>Analytics & optimization</strong><br>
Monitor campaign KPIs and customer engagement metrics. Generate insights and recommendations to improve conversion, retention, and customer experience. Support A/B testing, reporting dashboards, and marketing performance analysis.</p>
<p><strong>Leadership & collaboration</strong><br>
Collaborate with marketing, CRM, sales, analytics, and technology teams. Mentor junior team members and provide SFMC best practice guidance. Participate in solution architecture discussions and stakeholder workshops.</p>
<h4>Requirements</h4>
<ul>
<li>Bachelor’s degree in Computer Science, Information Systems, Marketing, or related field.</li>
<li>5+ years of hands-on experience with Salesforce Marketing Cloud.</li>
<li>Strong experience with Marketing Cloud Next capabilities and modern customer engagement platforms.</li>
<li>Proven expertise in:</li>
<ul>
<li>Journey Builder</li>
<li>Automation Studio</li>
<li>Email Studio</li>
<li>SQL</li>
<li>AMPscript</li>
<li>API integrations</li>
</ul>
<li>Experience integrating SFMC with Salesforce CRM, Data Cloud, and CDP solutions.</li>
<li>Strong understanding of customer lifecycle marketing and omnichannel strategies.</li>
<li>Excellent analytical, troubleshooting, and communication skills.</li>
</ul></p><p></p>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><b>Job Description </b><div> <p> <strong>The Opportunity: </strong> </p> <p>As a<strong> Part-Time Retail Merchandiser</strong>, you'll grow Mattel's presence at major retailers, working <strong>up to 30 hours per week</strong> (with extra hours during the exciting holiday season). We are looking for team members who are passionate about merchandising, enjoy interacting with customers & store management, and want to drive results!</p> <p> <strong>Why You ll Love It:</strong> </p> <p> 401K Matching & Bonus Opportunities<br>
Comprehensive Benefits<br>
Mileage Reimbursement & Paid Training<br>
Employee Referral Bonus<br>
Flexible Daytime Hours<br>
Growth Opportunities</p> <p> <strong>What Your Impact Will Be: </strong> </p> <ul> <li>Build strong relationships with store teams</li> <li>Brand ambassador of Mattel/FP Brands and products</li> <li>Secure prime product placements for Mattel s iconic brands</li> <li>Unpack cases of various Mattel toys</li> <li>Merchandise and organize product on the sales floor</li> <li>Construct & deconstruct store fixtures including shelving and racking</li> <li>Implement Point of Purchase materials</li> <li>Execute special projects as assigned including, but not limited to Store Events</li> <li>Complete reports as necessary including store surveys</li> <li>Train & lead team associates</li> <li>Provide real-time business insights</li> <li>Administrative tasks including but not limited to managing expenses, attending meetings, trainings, checking email etc.</li></ul></div></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"><b>Qualifications </b></p><div> <p> <strong>What We re Looking For: </strong> </p> <p> <strong>Basic Job Qualifications:</strong> </p> <ul> <li>Must be at least 18 years old</li> <li>High school diploma or GED, some college preferred</li> <li>Available to work daytime hours during the week; occasional required weekends</li> <li>Access to reliable internet access</li> <li>Must have reliable transportation to travel to between assigned store locations</li> <li>Reside within approximately <strong>20 miles of South Jordan, UT 84095</strong> or otherwise able to reliably service the assigned territory </li> </ul> <p> <strong>Physical Requirements:</strong> </p> <ul> <li>Stand, bend, twist, squat, kneel, reach</li> <li>Walk for long periods of time</li> <li>Climb ladders up to 12 tall</li> <li>Lift up to 25 pounds frequently, 50 pounds occasionally</li> <li>Push/pull carts containing product - occasionally up to 100 lbs.</li> <li>Operate power tools (screwdriver & drill)</li> </ul> <p> <strong>About You:</strong> </p> <ul> <li>Merchandising experience (preferred, but not required)</li> <li>Strong attention to detail and organizational skills, ensuring tasks are completed efficiently</li> <li>Develop strong relationships through effective communication and negotiation skills</li> <li>Represent the brand with professionalism and a positive attitude</li> <li>Ability to prioritize work with strong time management and a sense of urgency</li> <li>Self-motived with a strong bias for action by making resourceful decisions</li> <li>Eagerness to learn new skills with a safety-first mentality</li> <li>Tech-savvy (Microsoft Office, iPads and apps)</li> </ul> <p>Ready to make an impact? Apply today!</p> <p> <strong>The base hourly rate for this position is between $19.00 and $22.00.</strong> </p> <p>*The pay range is indicative of projected hiring range, however base pay will be determined based on a candidate s work location, skills and experience. Mattel offers competitive total pay programs, comprehensive benefits, and resources to help empower a culture where every employee can reach their full potential. </p></div><p></p></section>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $150/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply? Pass qualification(s)? Join a project? Complete tasks? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply? Pass qualification(s)? Join a project? Complete tasks? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply \u2192 Pass qualification(s) \u2192 Join a project \u2192 Complete tasks \u2192 Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply \u2192 Pass qualification(s) \u2192 Join a project \u2192 Complete tasks \u2192 Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>