ASP.NET Developers Jobs in Jordan
399 Jobs Found
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Place: Amman - Jordan Starting date: August 20th 2026 Duration of contract: 6 months Closing date for applications: 26th July 2026 Humanity & Inclusion is an independent and impartial aid organisation working in situations of poverty and exclusion, conflict and disaster.<br> The organisation works alongside people with disabilities and vulnerable populations, taking action and bearing witness in order to respond to their essential needs, improve their living conditions and promote respect for their dignity and fundamental rights.<br> At Handicap International-Humanity & Inclusion, we truly believe in the importance of inclusion and diversity within our organisation.<br> This is why we are engaged to a disability policy to encourage the inclusion and integration of people with disabilities.<br> Please indicate if you require any special accommodation, even at the first interview.<br> For more information about the organisation : https://hinside.<br>hi.org/intranet/jcms/a_16704/fr/accueil JOB CONTEXT: HI Palestine is currently operating as a stand-alone program under the authority of the Emergency Division, with a dedicated governance framework adapted to the scale and complexity of the crisis.<br> This arrangement, confirmed until the end of 2026, is intended to provide the structural flexibility and operational responsiveness required in the current context, while also anticipating and preparing for a transfer back to Mashreq Programme (regional office).<br> The mission's operational strategy relies on 4 pillars: Health (Rehabilitation and P&O), Armed Violence Reduction, Atlas Logistics, and Inclusive Education.<br> Cross-cutting components include IHA, MHPSS, Protection Mainstreaming, and DGA.<br> Women, men, girls and boys with disabilities in the West Bank and Gaza Strip have been disproportionately impacted by the recurrent escalations and unprecedented humanitarian crisis.<br> People with disabilities face significant needs and complex barriers to accessing humanitarian services and assistance across sectors of intervention.<br> YOUR MISSION MHPSS is a core component of HI’s intervention in Palestine.<br> MHPSS is systematically integrated into HI’s health activities, including rehabilitation and P&O, as well as into its education interventions.<br> In this context, the MHPSS Specialist plays a key role in ensuring the technical quality, coherence, and coordination of MHPSS activities across the programme.<br> As activities continue to expand in the WB, the position will also support the harmonization of approaches, tools, and processes between Gaza and the WB, ensuring consistency in programme implementation.<br> Finally, with the programme expected to transition to the regional office in the coming months, the MHPSS Specialist will contribute to preparing and supporting this transition, ensuring continuity of technical leadership and knowledge transfer.<br> Under the management of the Technical Head of Program (THoP) the MHPSS Specialist in Palestine (based in Amman, Jordan) continues the support and technical leadership for the MHPSS sector.<br> The MHPSS Specialist will ensure continuity and training and technically support MHPSS teams, while ensuring the quality, consistency, and integration of MHPSS across interventions and multidisciplinary health and education programmes.<br> The position will also support the harmonization of approaches between Gaza and the West Bank and contribute to preparing the programme's transition to the regional structure.<br> The MHPSS Specialist will have strong interactions with Internally: Area Manager, Project Managers, Technical officers, Officers, Specialists, MEAL department, Operations Manager, Grants Manager and externally: partners, local authorities, health structures, health NGOs.<br> The MHPSS Specialist will report to the THoP, will have a technical management from the MHPSS Global Specialist and will have the functionnal management of the MHPSS teams in Gaza and West Bank.<br> The position is open at both international and national levels.<br> If you are Jordanian, you will benefit from a national contract with the local package.<br> International contract: At HI, the conditions offered are up to your commitment and adapted to the context of your mission: 6 months International contract starting from August 20, 2026 based in Jordan; The international contract provides social cover adapted to your situation: Unemployment insurance benefits for EU nationals; Pension scheme; Medical coverage with 50% of employee contribution; Repatriation insurance paid by HI; Salary from 2757€ gross/month upon experience; Per diem: 640€ net/month - paid in the field; Hardship: No hardhsip allowance for Jordan Paid leaves: 25 days per year; R&R: 11 days/year Position: Unaccompanied Housing: Collective taken in charge by HI; YOUR PROFILE: At least least 5 years of overall professional experience in the field of MHPSS including a minimum of 3 years experience in emergency response in conflict affected countries.<br> You hold a Master degree in psychology, preferably in clinical psychology, or a related field is required.<br> Training in evidence-based therapeutic approaches for trauma, stress, and mental health issues in emergency contexts.<br> Comprehensive knowledge of the IASC Mental Health and Psychosocial Support in Emergency Settings guidelines and associated documents.<br> Experience in conducting multi-sectorial assessments is required Strong interpersonal and intercultural skills required; Fluency in oral and written English is compulsory; Arabic skills are a strong asset;</span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $150/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Optimiza is a regional systems integration and digital transformation company delivering consulting, business, and technology solutions across multiple industries.<br> Working at the intersection of business needs and technology execution, the company helps organizations modernize critical systems and deliver practical, scalable solutions that support long-term growth.<br> This role is suited to a developer who enjoys building and improving enterprise applications in complex environments.<br> As a Full-Stack Java Developer, you will contribute to initiatives that require strong technical judgment, close alignment with project goals, and the ability to support modernization efforts within business-critical systems.<br> Responsibilities Evaluate system performance and identify opportunities for improvement.<br> Implement reusable code and components across the application stack.<br> Follow up on project plans to ensure timely delivery and alignment.<br> Deploy delivered work, monitor outcomes, and troubleshoot issues.<br> Review and analyze business and technical requirements.<br> Class A Health Insurance Bachelor of Science in Computer Science, Engineering, or a related field.<br> 3-5 years of professional experience as a full-stack Java developer, including experience in core banking migration and legacy-to-modern core banking transformations.<br> Hands-on experience building Java EE applications with JSF and Spring MVC.<br> Strong foundation in object-oriented design and programming, with working knowledge of the JVM, including its limitations, weaknesses, and workarounds.<br> Experience implementing Web APIs and integrating services using Oracle stored procedures, TMF, and file-based interfaces.<br> Experience with relational databases, including MySQL and Microsoft SQL Server.<br> Track record of performing code reviews, writing unit and integration tests, and supporting continuous integration and test automation environments.<br> Experience contributing to architectural planning and refactoring.<br> Business-fluent English and Arabic.<br> Eligibility to work in Jordan.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $150/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<p><h4>Please submit your CV in English and indicate your level of English proficiency.<\/h4>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves:<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is NOT:<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for:<\/h4>\n<ul>\n <li>8+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard:<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p><\/p><p><\/p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<p><h4>Please submit your CV in English and indicate your level of English proficiency.<\/h4>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves:<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is NOT:<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for:<\/h4>\n<ul>\n <li>8+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard:<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 30 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p><\/p><p><\/p>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span><u><span><span>About the Job</span></span></u></span><p><span><span>As a </span></span><b><span><span>Senior Jasper Reports Developer</span></span></b><span><span> at Aspire, you will be responsible for designing, developing, and optimizing enterprise reporting solutions using Jaspersoft Studio (JasperReports). You will work closely with business stakeholders, database administrators, backend engineers, and QA teams to deliver scalable, high-performance reports that support business decision-making. This role requires strong expertise in SQL, Java, and report optimization, along with the ability to mentor junior developers and establish reporting best practices.</span></span></p><br><u><span><span>What you'll do</span></span></u><ul><li><p><span><span>Lead the design, development, and maintenance of complex enterprise reports using Jaspersoft Studio (JasperReports).</span></span></p><br></li><li><p><span><span>Design reusable report templates, sub reports, charts, cross tabs, and parameterized reports to support diverse business needs.</span></span></p><br></li><li><p><span><span>Develop and optimize complex SQL queries, stored procedures, and database objects for reporting solutions.</span></span></p><br></li><li><p><span><span>Integrate JasperReports with enterprise applications and APIs to deliver seamless reporting capabilities.</span></span></p><br></li><li><p><span><span>Gather, analyze, and translate business requirements into scalable and maintainable reporting solutions.</span></span></p><br></li><li><p><span><span>Optimize report performance for large datasets and high-volume environments.</span></span></p><br></li><li><p><span><span>Troubleshoot production issues, perform root cause analysis, and implement effective resolutions.</span></span></p><br></li><li><p><span><span>Ensure report quality, consistency, data accuracy, and security across reporting platforms.</span></span></p><br></li><li><p><span><span>Establish reporting development standards, coding guidelines, and best practices.</span></span></p><br></li><li><p><span><span>Conduct code reviews and mentor junior report developers to promote technical excellence.</span></span></p><br></li><li><p><span><span>Support report deployment, version control, and release management activities.</span></span></p><br></li><li><p><span><span>Collaborate with database administrators, backend developers, QA teams, and business stakeholders throughout the development lifecycle.</span></span></p><br></li><li><p><span><span>Prepare and maintain technical documentation, knowledge base articles, and implementation guides.</span></span></p><br></li></ul><u><span><span>What you'll need</span></span></u><ul><li><p><span><span>Bachelor's Degree in Computer Science, Information Technology, Software Engineering, or a related field.</span></span></p><br></li><li><p><span><span>Minimum of 5 years of hands-on experience developing enterprise reports using JasperReports/Jaspersoft Studio.</span></span></p><br></li><li><p><span><span>Strong expertise in SQL and relational databases such as Oracle, SQL Server, PostgreSQL, or MySQL.</span></span></p><br></li><li><p><span><span>Extensive experience with JRXML, report parameters, variables, expressions, sub reports, charts, and report optimization techniques.</span></span></p><br></li><li><p><span><span>Experience integrating JasperReports with Java-based applications using JasperReports APIs.</span></span></p><br></li><li><p><span><span>Strong knowledge of Java and object-oriented programming concepts.</span></span></p><br></li><li><p><span><span>Experience with Git and modern CI/CD pipelines.</span></span></p><br></li><li><p><span><span>Strong analytical, troubleshooting, and performance tuning skills.</span></span></p><br></li><li><p><span><span>Excellent communication and stakeholder management skills.</span></span></p><br></li><li><p><span><span>Experience working in Agile development environments is an advantage.</span></span></p><br></li><li><p><span><span>Familiarity with enterprise reporting architecture, data visualization, and reporting best practices is a plus.</span></span></p><br></li></ul><u><span><span>Why Aspire</span></span></u><p><span><span>In addition to a competitive long-term total compensation package with salary and performance-based bonus, we have a reward philosophy that goes beyond compensation.</span></span></p><br><ul><li><p><span><span>Be part of a remote-first organization where flexibility is embraced.</span></span></p><br></li><li><p><span><span>Work and learn alongside talented engineers and technology leaders.</span></span></p><br></li><li><p><span><span>Explore opportunities to learn and grow through technical and non-technical training programs.</span></span></p><br></li><li><p><span><span>Gain global exposure by working on products with international teams and clients.</span></span></p><br></li><li><p><span><span>Enjoy our nursery reimbursement benefit.</span></span></p><br></li><li><p><span><span>Attend virtual and in-person international technology conferences to expand your knowledge and network.</span></span></p><br></li></ul><br> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> About the Role You’ll design coding tasks that challenge frontier AI coding agents.<br> Each task is a self-contained Docker environment with a broken piece of software; an AI agent attempts the fix; automated tests verify the outcome.<br> Your deliverable is the full task package: broken code, tests, instructions, and a reference solution proving the task is solvable.<br> Responsibilities : Invent a realistic developer scenario — a real bug, a broken ETL, a missing feature — not a toy problem.<br> Build a reproducible Docker environment with pinned dependencies.<br> Write a pytest that verifies outcomes, not specific commands — deterministic, non-flaky, and does not leak the fix.<br> Write an instruction.<br>md that reads like a Jira ticket a developer would receive.<br> Write a reference solve.<br>sh proving the task is solvable.<br> Calibrate difficulty so current state-of-the-art agents solve the task 20–60% of the time.<br> Iterate based on feedback from expert QA reviewers.<br> Later: review other authors’ tasks as a QA reviewer.<br> Not in scope Data labeling, prompt engineering.<br> Production code to ship — you design problems and verification for AI agents.<br> Leetcode puzzles — scenarios must look like real developer work.<br> Not every candidate task ships — quality over quantity.<br> Requirements 3+ years of production software development in one backend stack — Python, Go, Node.<br>js, Java, or Rust.<br> Depth in one stack beats breadth.<br> Python + pytest fluency — required regardless of primary stack.<br> The task harness is pytest-based even when the broken app is in another language.<br> Fixtures, parametrize, monkeypatch, timeouts, conftest.<br>py. Docker authoring — reproducible Dockerfiles, pinned dependencies, multi-stage builds when needed, non-root user.<br> Linux & Bash — comfort debugging inside containers (strace, lsof, journalctl); shell beyond set -euo pipefail.<br> AI coding agent experience — Claude Code, Cursor, Roo Code, or similar, on non-trivial work.<br> You can cite a specific time the AI was confidently wrong and how you caught it.<br> English — B2+ written.<br> Not a fit Data Science, ML, or Computer Vision engineers without backend-engineering output.<br> Manual QA testers without automation or test authoring.<br> Frontend-only, low-code / no-code, IT Support, or Business Analysts.<br> Engineers who have never written pytest from scratch.<br> Junior, intern, or assistant as the most recent role.<br> Preferred qualifications Domain depth in Security, System Administration (nginx / systemd / cron), Scientific Computing (NumPy / PyTorch / SciPy), DevOps, or Git internals.<br> Modern Python tooling (uv, poetry, pyproject.<br>toml). Coverage tooling (pytest-cov, coverage.<br>py, gcov, llvm-cov, kcov).<br> Fuzzing or property-based testing (Hypothesis).<br> Prior contribution to agent-evaluation benchmarks or related frameworks.<br> Process Apply → Pass qualification (90-minute sample-task screen + short behavioral interview) → Join a project → Complete tasks → Get paid.<br> Time commitment Onboarding: ~10 hours per first task.<br> Steady state: ~5 hours per task, 2–4 parallel tasks per author.<br> Realistic weekly load: 8–20 hours.<br> Higher volume available for top performers.<br> You choose when and how to contribute; tasks must be submitted by the deadline and meet acceptance criteria.<br> Compensation: Paid contributions, rates up to $35/hour *.<br> Task-based compensation equivalent to hourly rate, depending on performance and volume.<br> Some projects include incentive payments.<br> *Rates vary based on expertise, skills assessment, location, project needs, and other factors.<br> Higher rates may be provided to highly specialized experts.<br> Lower rates may apply during onboarding or non-core project phases.<br> Payment details are shared per project.<br> Apply Submit your CV via the Mindrift platform.<br> Indicate your English level, note this role (Software Engineering Evaluation Specialist — Terminal Bench), and include a GitHub profile link if available.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br>What this opportunity involves: We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT: Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for: 8+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard: Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paidEffort estimate Tasks for this project are estimated to take 30 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation: Up to $150/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~30 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>Please submit your CV in English and indicate your level of English proficiency.<br> Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems.<br> Participation is project-based, not permanent employment.<br> What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<br> You'll create challenging tasks and evaluation criteria within realistic simulated environments: Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust What this is NOT Not data labeling Not prompt engineering Not writing code from scratch - the agent writes most of the code; you guide and evaluate What we look for 5+ years in software development Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis Experience writing tests (functional, integration) English proficiency - B2+ Why this is hard Frontier models are already good at coding.<br> Creating a task that genuinely challenges the best models is non-trivial.<br> You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution.<br> Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<br> How it works Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid Effort estimate Tasks for this project are estimated to take 20 hours to complete, depending on complexity.<br> This is an estimate and not a schedule requirement; you choose when and how to work.<br> Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<br> Compensation Up to $50/hr equivalent , depending on level and pace.<br> Tasks are estimated at ~20 hours each; you set your own schedule.<br></span> </div>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n <li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n <li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n <li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n <li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n <li>Not data labeling<\/li>\n <li>Not prompt engineering<\/li>\n <li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n <li>5+ years in software development<\/li>\n <li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n <li>Experience writing tests (functional, integration)<\/li>\n <li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply? Pass qualification(s)? Join a project? Complete tasks? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at approximately 20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>
<p><h4>Description<\/h4>\n<p>Please submit your CV in English and indicate your level of English proficiency.<\/p>\n<p>Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.<\/p>\n<h4>What this opportunity involves<\/h4>\n<p>We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.<\/p>\n<p>You'll create challenging tasks and evaluation criteria within realistic simulated environments:<\/p>\n<ul>\n<li>Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history<\/li>\n<li>Design tasks from intermediate states of these environments - craft the prompt, define what \"solved\" means, and ensure the task is solvable by an AI agent<\/li>\n<li>Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient<\/li>\n<li>Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust<\/li>\n<\/ul>\n<h4>What this is not<\/h4>\n<ul>\n<li>Not data labeling<\/li>\n<li>Not prompt engineering<\/li>\n<li>Not writing code from scratch - the agent writes most of the code; you guide and evaluate<\/li>\n<\/ul>\n<h4>What we look for<\/h4>\n<ul>\n<li>5+ years in software development<\/li>\n<li>Core stack: Python (FastAPI), JavaScript\/TypeScript (React), Docker, Postgres, Kafka, Redis<\/li>\n<li>Experience writing tests (functional, integration)<\/li>\n<li>English proficiency - B2+<\/li>\n<\/ul>\n<h4>Why this is hard<\/h4>\n<p>Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.<\/p>\n<h4>How it works<\/h4>\n<p>Apply ? Pass qualification(s) ? Join a project ? Complete tasks ? Get paid<\/p>\n<h4>Effort estimate<\/h4>\n<p>Tasks for this project are estimated to take 20 hours to complete, depending on complexity. This is an estimate and not a schedule requirement; you choose when and how to work. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.<\/p>\n<h4>Compensation<\/h4>\n<p>Up to $50\/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.<\/p><\/p><p><\/p>