Senior Software Engineer, AI Coding Evaluation
- $80 - $100 per hour
- $1,600 - $2,000 for 20 hours
- Remote, any location
We're looking for experienced software engineers to build one AI coding evaluation task each. You'll work from a codebase you know well. The task takes 15 to 20 hours and is due Tuesday 29 September.
About the role
You pick a real change in a repository you know from the inside, one that a leading AI coding model gets wrong. You write the fix, the tests that prove it, and a short grading rubric.
The task is used to measure and train AI coding models.
You start with a one-page proposal. We approve it before you build anything.
What you'll do
- Proposal
- A one-page proposal for the change you'll build, approved by us before you start.
- Environment
- The repository building and running its tests offline in Docker, at a pinned commit.
- Your fix
- Your own fix, written to the standard a maintainer would merge.
- Blocker tests
- Tests that reject a nearly-right fix.
- Review criteria
- 3 to 5 review criteria covering style, scope and design.
- Model runs
- 3 runs of the AI model on your task, showing where it fails. The runs use a free tier, so they cost you nothing.
Requirements
- Experience
- 5+ years of professional software engineering.
- Open source
- Code merged into a public open-source repository, ideally one you know well. A private codebase you own also works. Employer or client code doesn't.
- Languages
- Strong in at least one of Python, TypeScript, Go, Rust, Java or C++.
- Tooling
- Solid Docker, Linux, Git and testing skills.
- AI tools
- You use an AI coding tool regularly.
- AI evaluation
- Experience writing tests, rubrics or evaluation criteria for AI models.
- Availability
- Available now for 15 to 20 hours. The task is due Tuesday 29 September.
- Device security
- Work happens on a secure personal or work computer. Public or shared machines aren't permitted.
Nice to have
Not required. Worth mentioning if it applies to you.
- Deep expertise in one area, such as databases, compilers, distributed systems, security or infrastructure.
- Experience building agent tasks or RL environments.
Contract and payment
- You sign Askable's contract and accept its terms before you start.
- Fully remote, from any location.
- You're paid only for a task that passes our quality review. We pay you directly, not through an agency.
- The work will be used for AI training. The change and your solution must be your own, and you must be free to license them to us.
- Don't publish the work anywhere. That means no public pull request and no public fork.
About Askable
Askable is the research infrastructure powering global enterprises since 2017. Askable Labs, a division of Askable Pty Limited, produces primary data for frontier AI labs from experts doing real work. We recruit those experts ourselves, verify them, and pay them directly.
We consider every qualified applicant without regard to legally protected characteristics, and we make reasonable adjustments on request.