Cortext connects experts with the world's top AI labs. Rate answers, write code, teach languages, review research — on your own schedule. Get paid weekly, in USD.
Training data and evaluation for the labs building these models
Marks shown belong to their respective owners. Their appearance here describes the field Cortext works in and does not indicate a commercial relationship, partnership or endorsement.
Across 18 domains, from entry annotation up to $250/hr for credentialed clinical, legal and quantitative work. Every listing publishes its rate band, its pay unit and its expected weekly hours before you apply — and the band does not change with the country you live in.
Four things happen between submitting an application and money arriving. Here is what each one costs you in time, and what it pays.
One application covers the whole board. Cortext verifies your credential, you sit one domain assessment, and after that you are matched to projects without reapplying. Of 214,300 applications in 2025, 61% cleared credential review and 34% passed assessment.
Read the full process →The pass mark on the domain assessment is 80%, and a failed attempt cannot be retaken for 30 days. We would rather tell you that now than after you have spent an evening on it.
What the screen covers →The assessment runs 60–90 minutes and a paid trial task runs 3–5 hours. Anything beyond the first unpaid 30 minutes is billable at the role's advertised rate, including the trial you fail.
Screening detail →Median time from admission to a first paid task is 17 days. Every listing publishes its rate band, its unit of pay and its expected weekly hours before you apply to it.
See open listings →The pay week runs Monday 00:00 UTC to Sunday 23:59 UTC, review takes 3 days, and the payout run goes out Friday, 12:00 UTC over Stripe and Wise. Minimum payout is $10; standard transfers land in 2–5 business days. Your first payout is held 7 days for fraud screening — after that it is weekly without exception.
Payment mechanics →Paid once your referral bills 20 hours. Credentialed seats that are genuinely hard to fill pay up to $900. Bonus runs go out Tuesday and Friday, 12:00 UTC.
Referral terms →Counts and rate bands below are read off the live board, not written by hand. If a domain is not listed, it has nothing open today.
Translate, post-edit and evaluate model output in your native language, and write the reference answers a lab grades its system against.
Review model-written code against a compiling test harness, author reference solutions, and explain in writing why a rejected patch is wrong.
Apply a lab's rubric to paired model responses, record the deciding criterion, and flag rubric items that cannot be applied consistently.
Review clinical reasoning against current guidance, correct unsafe generations, and write the citation that supports the correction. Licence verified before your first task.
Label, span-tag and adjudicate against a published guideline, including the disagreement cases two earlier annotators could not settle.
Author and verify graduate-level problems and solutions, and grade proofs step by step rather than on the final answer.
Annotate multi-step tool-use trajectories: mark the step where the agent went wrong, classify the failure, and write the correction the model should have produced.
Check quantitative and accounting reasoning, reproduce the calculation, and record where a model's method diverges from standard practice.
Assess statutory and contractual reasoning within a named jurisdiction, and mark the point at which an answer stops being defensible.
Evaluate long-form and stylistic generation against a brief, and rewrite the passages that miss it.
Probe deployed systems for policy failures, document reproducible attack paths, and rate severity against the lab's own harm taxonomy.
Grade image and video understanding, caption to specification, and mark where a model has described something that is not in the frame.
Transcribe, diarise and score speech output for accent, disfluency and prosody handling, including low-resource and code-switched audio.
Verify query, pipeline and statistical output, and mark analyses that are technically valid but answer a different question.
Grade tutoring and explanation quality at a stated level, and rewrite explanations that are correct but unteachable.
Grade coding, prior-authorisation and documentation workflows, and mark output that would fail a payer audit or a compliance review.
Structure and grade business-analysis output against the evidence actually supplied in the prompt.
Check model advice against installed practice and code, and document where a plausible answer would fail inspection.
Multiply the published bands by the hours you would realistically give it. The output is a range, because the input is a range.
Drawn from the 20 language listings currently priced per hour, whose published bands run $34–$105 an hour. Where you land inside that depends on the credential the listing requires and your review scores, not on where you live.
An estimate of gross contract income, not a guarantee and not an offer. It assumes you are admitted to this domain and that work is available every week, which is not how the queue behaves — volume tracks what the labs are training that quarter and can thin out with little notice. Nothing here is deducted for tax; contractors are responsible for their own. How payment works →
Six people on six different queues, including the parts they did not enjoy.
I take clinic Monday to Thursday and grade oncology reasoning on Friday mornings. The part I did not expect was being asked to cite the guideline behind every correction — it is closer to peer review than to survey work, and it is the reason I have kept doing it for fourteen months.
My first trial task came back rejected with three paragraphs explaining which of my test cases did not actually exercise the bug. Annoying, and then useful. I passed the retake and I have not had an hour removed since.
Yoruba work used to mean being offered a third of what the Spanish queue paid for identical instructions. Here the band was published before I applied and it was the same band. I check, because I used to have to.
Agent trajectories are the only annotation work I have found that is genuinely hard. You have to decide which of eleven steps was the actual mistake, and the rubric will not let you hand-wave it. Two hours in I am tired in a way that survey labelling never managed.
The weekly payout is the whole thing for me. I know what lands on Friday and I know the three-day review window that produced it. I have been on platforms where an invoice sat for sixty days and nobody would say why.
Red-team work dried up for six weeks when the project I was on finished, and Cortext said so directly instead of pretending the queue was temporarily empty. I would rather plan around that than be strung along.
Quotes are lightly edited for length. Earnings and hours described are individual and are not a projection of what anyone else will make.
The ones that decide whether this is worth your evening. Longer answers, including tax forms and the appeals process, are on the support and how-it-works pages.
Full support centre →Weekly payout run at Friday, 12:00 UTC over Stripe and Wise, $10 minimum, no invoicing.
No signup fee, no platform cut, no withdrawal fee, and the FX spread is absorbed rather than passed on to you.
Every listing shows its band, its pay unit and its expected weekly hours on the card. Rates do not vary by country.
Rejections name the rubric item that failed and can be appealed to a second reviewer.
Apply once, get matched to projects that fit your skills, and cash out every week.
Find your role →