Explore opportunities
Software Engineer Expert — Multi-Language Review
IntermediateAuthor repo-scale coding tasks in Python, Go or Java and adjudicate the model's patch against your own tests.
Frontend Engineer — Code Evaluation (React)
SeniorGrade model-written React on the things tests miss: accessibility, re-render cost and state that should never have been local.
Distributed Systems Expert — Failure Reasoning
SeniorWrite scenarios where consensus, clock skew or partial failure breaks the system, and score whether the model reasons or guesses.
CUDA & GPU Kernel Engineering Expert
SeniorBenchmark model-written kernels for occupancy, bank conflicts and numerical drift, then explain why they are slow.
MLOps Engineer — Training Stack Evaluation (JAX / PyTorch)
SeniorJudge whether model-proposed sharding, checkpointing and data-loading advice would actually survive a multi-node run.
Site Reliability Engineer — Incident Trace Evaluation
SeniorReplay real incident timelines and score whether the model's next action would have helped, done nothing, or made it worse.
Legacy Systems Expert — COBOL, ABAP & Mainframe
SeniorVerify model reasoning about COBOL copybooks, JCL and ABAP that no public training corpus covers well.
Embedded & Firmware Engineering Expert
SeniorScore model answers on interrupt safety, DMA coherency and timing budgets where a plausible answer bricks the device.
Database Internals & Query Optimization Expert
SeniorGrade model query rewrites against real execution plans, isolation semantics and index cost — not against intuition.
Network Engineer — Packet & Telemetry Reasoning
SeniorHand the model a packet capture and score whether it finds the retransmit storm or invents a plausible cause.
QA Engineer — Model Output Verification
IntermediateTurn model-written test suites inside out: do they actually fail when the code is broken?
Agentic Trajectory Evaluator — Tool-Use Traces
IntermediateRead an agent's whole run — every tool call, every retry — and mark the step where it stopped making sense.
Computer-Use Annotation — Desktop & Browser Agents
EntryRecord yourself doing ordinary computer chores, one deliberate click at a time, with a reason typed for each.
Agentic Reward Model Design — Process Supervision
SeniorDesign the scoring function that decides which agent behaviours get reinforced, then find how it will be gamed.
Web Agent Grounding — DOM Reference & Element Resolution
IntermediateDecide whether the agent clicked the right thing for the right reason, or the right thing by accident.
Tool & API Schema Authoring for Agent Harnesses
IntermediateWrite the tool definitions agents are given, including the badly-written ones they have to survive.
Long-Horizon Agent Review — Plans Over 100+ Steps
SeniorJudge whether an agent that ran for three hours was making progress or elaborately standing still.
RLHF Preference Ranking — General Assistant
IntermediateChoose between two answers that are both fine, and write down the reason that would hold up if someone disagreed.
Rubric Grading — Long-Form Reasoning Output
IntermediateScore multi-thousand-word answers against a six-dimension rubric without letting one bad paragraph sink the rest.
LLM-as-Judge Validation — Human Reference Set
SeniorFind the cases where the automated judge and a careful human diverge — that gap is the deliverable.
Factuality & Citation Grounding Reviewer
IntermediateCheck every citation the model gave you and record which ones support the sentence they are attached to.
Instruction-Following & Constraint Adherence Evaluator
EntryCount the constraints in the prompt, check each one, and ignore how good the answer otherwise is.
SFT Gold Response Author — General Knowledge
IntermediateWrite the answer the model should have given — the reference, not a correction of what it produced.
DPO Pair Construction — Targeted Contrast Sets
SeniorBuild chosen/rejected pairs that differ on exactly one thing, so the gradient has nowhere else to go.
Multi-Turn Conversation Evaluator — Context Retention
IntermediateRun long conversations designed to break the model's memory, then mark the exact turn it forgot something.
Text Classification — Taxonomy Application at Scale
EntryApply a 60-label taxonomy to short text, and be the person who notices when two labels overlap.
Named Entity & Span Annotation — Technical Corpora
IntermediateDecide where an entity actually starts and ends, in text where the answer is genuinely contested.
Conversational Data Collection — Scripted Personas
EntryPlay a specific person with a specific problem, convincingly enough that the transcript is worth training on.
Search & Retrieval Relevance Rating
EntryJudge whether a retrieved passage answers the query or merely shares its vocabulary.
Annotation Guideline Author & Quality Lead
SeniorWrite the document fifty annotators will read literally, then fix every place they read it differently.
Sentiment & Stance Annotation — Irony and Understatement
IntermediateLabel the text where positive words carry negative meaning, and say which cue told you.
Document Extraction Annotation — Forms, Tables & Invoices
EntryPull the right number out of a badly scanned table, and record how you knew which column it was in.
Vision-Language Grounding — Caption Verification
IntermediateCheck whether the model is describing the image in front of it or the image it expected to see.
Chart, Table & Document VQA Authoring
IntermediateWrite questions about charts that cannot be answered by reading one number off an axis.
Video Temporal Annotation — Action & Event Boundaries
EntryMark the frame where an action begins, and argue for it when a colleague picks a different one.
Handwriting & Historical Document Transcription
IntermediateTranscribe what the writer actually wrote, including the misspelling, and mark what you genuinely cannot read.
Multimodal Output Review — Image Generation Quality
IntermediateScore generated images on prompt fidelity, and separately on the anatomy nobody wants to look at closely.
ASR Transcription & Correction — Accented Speech
IntermediateCorrect machine transcripts of speech the machine was never trained to hear, and log why it failed.
Speech Synthesis Rating — Naturalness & Prosody
EntryListen to synthetic speech and say precisely what is wrong with it, not just that it sounds off.
Voice Data Collection — Scripted & Spontaneous Speech
EntryRecord your own voice reading prompts and talking naturally, in your own accent, deliberately unpolished.
Audio Event & Speaker Diarisation Annotation
IntermediateMark who is speaking and when, especially in the four seconds where three people are talking at once.
Swahili Gold Response Author — Assistant Fine-Tuning
IntermediateWrite the Swahili answers the model should give, in Swahili it would not produce by translating from English.
Gĩkũyũ Low-Resource Corpus Builder
IntermediateBuild first-party Gĩkũyũ text where almost none exists, with tone marked and provenance recorded.
Dholuo–English Code-Switching Data Collection
IntermediateProduce and annotate the mid-sentence language switching that real speakers do and corpora never capture.
Hindi Fluency & Cultural Accuracy Rater
EntryRate Hindi output on whether a Hindi speaker would say it, separately from whether it is grammatical.
Japanese Register & Keigo Reviewer
IntermediateCheck whether the honorific level matches the relationship, including the times it is too polite.
Portuguese Translation Post-Editing — pt-BR and pt-PT
IntermediateFix machine translation into Portuguese and log which errors the engine keeps making.
Yorùbá Corpus Builder — Tone Marking & Diacritics
IntermediateWrite and correct Yorùbá with full diacritics, in a corpus where almost everything online has none.
Vietnamese Gold Response Author — Technical & Everyday
IntermediateAuthor Vietnamese reference answers where the pronoun choice alone encodes who is speaking to whom.
Tagalog–English Code-Switching & Taglish Data
IntermediateProduce Taglish as it is actually spoken and annotate what triggers each switch mid-clause.
Korean Honorific & Speech-Level Reviewer
IntermediateCheck the speech level holds for a whole message, because switching mid-paragraph is the tell.
German Technical Register Reviewer — Compounds & Formality
IntermediateReview German output for invented compounds, Anglicisms and the Sie/du choice the model gets wrong.
French Post-Editing & Terminology — Metropolitan and Canadian
IntermediateBring machine French to publication standard and classify every edit so the engine's profile is measurable.
Bengali Gold Response Author — Bangladesh & West Bengal
IntermediateWrite Bengali reference answers in one variety on purpose, rather than the blend the model produces.
Urdu Reviewer — Nastaʿlīq Orthography & Register
IntermediateReview Urdu for the orthography and vocabulary that separate it from Hindi in everything but grammar.
Amharic Corpus Builder — Fidäl Script & Technical Register
IntermediateWrite technical Amharic that mostly does not exist yet, and record the terms you had to settle.
isiZulu Fluency & Noun-Class Reviewer
IntermediateCheck the concord agreement chain, which is where model isiZulu falls apart in longer sentences.
Hausa Corpus & Hausa–English Code-Switching
IntermediateProduce written Hausa in Boko script with tone and length marked, plus the switching real speakers use.
Somali Low-Resource Corpus Builder
IntermediateWrite original Somali across domains the digital corpus simply does not have, with dialect recorded.
Thai Reviewer — Segmentation, Register & Particles
IntermediateReview Thai where the model got the words right and the word boundaries, particles or politeness wrong.
Polish Gold Response Author — Case and Aspect Discipline
IntermediateWrite Polish reference answers where case government and verbal aspect are chosen, not guessed.
Turkish Post-Editing — Agglutination & Vowel Harmony
IntermediatePost-edit Turkish where a single wrong suffix in a seven-suffix chain changes who did what to whom.
Adversarial Red Teamer — Jailbreaks & Prompt Injection
SeniorBreak the model's refusal behaviour on purpose, document the technique, and check whether the patch actually held.
Harm Taxonomy Annotation — Policy Boundary Cases
IntermediateLabel the requests that sit exactly on the policy line, where both refusing and answering are defensible.
Biosecurity Evaluation — Uplift Assessment (PhD)
PhDJudge whether a model's answer would meaningfully help someone who should not be helped, using expertise almost no one has.
Offensive Security Evaluation — Cyber Capability Assessment
SeniorMeasure how far a model can actually get on an intrusion chain, in a range where it cannot reach anything real.
Content Policy Adjudicator — Escalation Tier
SeniorDecide the cases two trained reviewers could not agree on, and write the rule so the next one is easier.
Physician Clinical Reasoning Evaluator (MD / DO)
MDGrade the differential, not the final answer — most model errors are a correct diagnosis reached by an unsafe route.
Radiologist — Imaging Report Evaluation (Board-Certified)
MDRead the study yourself, then judge whether the model's report would be safe to sign.
Oncologist — Treatment Reasoning Gold Author
MDWrite the reference answer for cases where the guideline runs out and the tumour board actually decides.
Psychiatrist — Mental Health Safety Reviewer
MDJudge risk responses on both failure directions: missing the risk, and treating every low mood as an emergency.
Pharmacist — Medication Safety & Interaction Review (PharmD)
Master'sCheck dosing, interactions and the renal adjustment the model silently left out.
Registered Nurse — Patient Communication Reviewer
IntermediateJudge whether a patient could actually follow the instruction, and whether they would still ask the real question.
Clinical Documentation Annotation — De-Identified Notes
IntermediateAnnotate what a clinical note actually asserts, including the findings it explicitly rules out.
Guideline Concordance Reviewer (MD or PhD Epidemiology)
PhDCheck the model's citations against the actual evidence, including whether the trial population matches the patient.
Medical Coding & Claims Reviewer (CPC / CCS)
IntermediateJudge the model's code assignment against the documentation, including the specificity it invented.
Prior Authorization & Utilization Review Specialist
IntermediateDecide whether the criteria were actually met, and whether the model's denial would survive appeal.
Revenue Cycle & Denials Workflow Reviewer
IntermediateEvaluate whether the model's denial-resolution plan would actually get the claim paid.
Clinic Operations & Scheduling Workflow Annotator
EntryAnnotate the front-desk decisions that look trivial and hold a clinic's day together.
Commercial Contract Review Evaluator (JD, Bar-Admitted)
JDThe model reads each clause correctly and loses the obligation the cross-reference three pages later imposes.
Litigation Reasoning & Motion Practice Reviewer (JD)
JDCheck every citation to the reporter, then check whether the case still says what the model claims it says.
Privacy Counsel — Cross-Regime Compliance Reviewer
JDAnswers that are correct under GDPR and quietly wrong under the three other regimes the same product operates in.
Patent Practitioner — Prior Art & Claim Construction Reviewer
JDClaim scope is decided by one word in the preamble, and the model reads the abstract.
Litigation Support Paralegal — Document Review Annotation
IntermediateLabel responsiveness, privilege and issue tags on discovery documents where the hard calls are all partial.
Consumer Legal Communication Reviewer (Bar-Admitted)
JDEvery plain-language rewrite is a redraft, and a redraft can change what the notice legally does.
Quantitative Analyst — Derivatives Pricing Reasoning (PhD)
PhDThe derivation is clean, the model is the wrong one for the payoff, and nothing in the output says so.
Equity Research Analyst — Valuation Model Reviewer (CFA)
SeniorA DCF whose arithmetic is perfect and whose terminal growth rate quietly exceeds long-run GDP.
Risk Model Reviewer — Market & Credit (FRM)
SeniorRisk numbers computed correctly under an assumption the scenario was specifically designed to break.
Technical Accounting Reviewer — IFRS & US GAAP (CPA / ACA)
SeniorAnswers that blend IFRS and US GAAP into a single set of rules that exists in neither.
AML & Sanctions Alert Review Annotator (CAMS)
IntermediateLabel why an alert closes, not just that it closes — the reason is the only part a regulator reads.
Corporate Tax Reviewer — Cross-Border Structuring (CPA / CTA)
SeniorTreaty analysis that is right on the article and wrong on which country's domestic law gets there first.
Management Consultant — Strategy Case Reasoning Evaluator
SeniorA structure that partitions the problem beautifully and never gets to a number anyone could act on.
Operations & Supply Chain Consultant — Diagnostic Reviewer
SeniorDiagnoses that name the bottleneck the data shows and miss the one the plant floor would have told you about.
K–12 Teacher — Instructional Response Evaluator (Registered)
IntermediateExplanations that are correct, age-appropriate and give away the answer the student was about to reach.
Assessment Designer — Item Writing & Psychometric Review
SeniorItems that test whether the student read carefully rather than whether they understood anything.
Special Education Specialist — Accommodation Content Reviewer
SeniorAccommodations that are legally correct, generically stated, and impossible for a teacher to implement on Tuesday.
University Instructor — Essay Grading Calibration
Master'sGrades that track how well an essay is written rather than whether its argument holds.
Mathematician — Proof Verification & Gap Detection (PhD)
PhDFind the step that is asserted rather than proved, and say what would have to be true for it to hold.
Physicist — Problem Authoring & Solution Grading (PhD)
PhDSolutions that reach the right number with units that do not survive a dimensional check.
Synthetic Chemist — Reaction Route Evaluator (PhD)
PhDRoutes that are textbook-correct on paper and would fail at step three in an actual fume hood.
Molecular Biologist — Experimental Design Reviewer (PhD)
PhDProtocols with every reagent and no control that would tell you the result meant anything.
Statistician — Study Design & Inference Reviewer (PhD)
PhDThe test is applied correctly, the assumption it needs was never checked, and the p-value is reported to three decimals.
Mechanical Engineer — Design Calculation Reviewer (PE / CEng)
SeniorCalculations that pass on the load case that was given and were never checked against the one that governs.
Power Systems Engineer — Protection & Load Study Reviewer
SeniorCoordination curves that grade correctly at fault current and lose selectivity everywhere else on the curve.
Fiction Writer — Narrative Craft Evaluator
SeniorProse with no bad sentences in it and no reason for anyone to read the second page.
Screenwriter — Scene & Dialogue Reference Author
SeniorWrite the reference scenes where the characters say the wrong thing and mean the right one.
Brand Copywriter — Voice Consistency Reviewer
IntermediateCopy that follows every rule in the brand guide and sounds like a different company each time.
Art Director — Generated Image Critique
SeniorTechnically flawless images that solve none of the problem the brief actually set.
Composer — Generated Music Structure Evaluator
SeniorEight bars that work, repeated until the piece ends, with nothing developing in between.
Licensed Electrician — Code Compliance Reviewer
IntermediateAnswers that cite the right article of a code cycle the jurisdiction has not adopted yet.
HVAC Technician — Diagnostic Reasoning Annotator
IntermediateDiagnostic trees that reach the failed part by replacing the four cheaper ones first.
Master Automotive Technician — Fault Diagnosis Evaluator (ASE)
IntermediateReading a fault code as a diagnosis rather than as a symptom, which is how good parts get replaced.
Analytics Engineer — SQL & Metric Definition Reviewer
IntermediateQueries that run, return a number, and answer a slightly different question than the one asked.
Experimentation Analyst — A/B Test Reasoning Evaluator
SeniorReading a result as a decision when the experiment could never have detected the effect that mattered.
BI Practitioner — Dashboard & Visualisation Critique
IntermediateCharts that are technically valid and mislead anyone who reads them at a glance, which is everyone.
Applied ML Practitioner — Model Failure Diagnosis Reviewer
SeniorDiagnoses that tune the hyperparameters of a model whose training set leaked the label.