About rustLabs
Expert humans in the training loop.
rustLabs hand-builds RL environments, rubrics, and expert evaluations for the labs training frontier models — recruited, coordinated, and quality-checked through the rustBench network.
What we build
Human-in-the-loop work for evals, post-training, and agents.
RL environment design
Interactive tasks where your model acts, gets feedback, and is graded on the outcome — not just the final token.
Reward & rubric writing
We turn messy domain standards into measurable pass/fail, partial-credit, and preference rules a trainer can use.
Agent trace review
Experts inspect tool calls, reasoning paths, edits, and citations to find exactly where a run goes wrong.
Expert evaluation
Field specialists compare model outputs against real acceptance criteria, with written rationale — not vibes.
Factual verification
Claims, calculations, code behavior, and policy checked against ground truth by people who know the domain.
Benchmark & eval tasks
Prompts, repos, and adversarial cases with known-good outcomes, ready to gate your next release.