From research objective to reliable data and evaluation.
Owen Onderdonk · AI Data & Evaluation Program Lead
An evaluator can be useful without being entitled to decide.
AutoQA Foundation is my independent public framework for the capability and authority boundaries of LLM-based quality review. It begins with a simple constraint: the record sets the ceiling. A model cannot recover facts absent from the observable record, make an ambiguous instruction decide itself, or earn authority by sounding confident.
Record identifiability
When the same observable record can correspond to different correct decisions, no evaluator that sees only that record can be correct in every case. Better processing does not manufacture missing facts.
- Task boundary
- Capability boundary
- Authority boundary
Four honest outcomes
Uncertainty is routed, not hidden. A binary pass/fail interface does not remove ambiguity from the record -- it converts it into confident error.
- Supported
- Contradicted
- Insufficient
- Contract-ambiguous
Authority must be earned
A capable detector is not automatically entitled to impose consequences. Authority is granted lane by lane, on measured evidence, with a named human owner.
- Lane-level calibration
- Measured risk
- Named human ownership
✦ Independent public working framework -- not a certification body, formal standard, or claim that a particular automated reviewer is accurate.
Selected work has included managing data and evaluation projects supporting model development for Anthropic, Google, and Amazon through external AI data partners.
These engagements are complete; project details remain confidential.
I Learned the System From the Inside Out
I entered frontier-model data work in February 2024, creating difficult mathematics and physics training data. I moved into reviewing contributors, then auditing the decisions of other reviewers -- and from there into designing the standards, calibration processes, and workflows that govern the work.
In February 2025 my scope expanded into project development, program management, client collaboration, and quality leadership. I helped develop and run programs with a workforce of more than 3,000 contributors -- work that supported model development for Anthropic, Google, and Amazon through external AI data partners -- and built one emerging-capability evaluation framework through seven prototypes over three months of client collaboration.
That progression is the point: each layer -- contributor, reviewer, reviewer of reviewers, systems designer, program lead -- changed what I could see. Learning Voyage is my independent practice and the home for this work as I extend it into specialized-data partnerships and proprietary-data systems.
Owen Onderdonk
AI Data & Evaluation Program Lead · Founder, Learning Voyage LLC
- Helped develop and run AI data programs with a 3,000+ contributor workforce
- Built an emerging-capability evaluation framework through seven prototypes
- Hands-on evaluation of training data and AI-generated code for frontier models
- 15 years designing STEM and technical education programs
- B.S. in Physics · Master's degree, Randolph College
The Layer Between a Hard Question and Dependable Evidence.
I work where research intent, domain expertise, data production, human judgment, and operational delivery have to become one coherent system. Each area is labeled so established work is never confused with direction or development.
Research data operations Established
Translate capability goals and failure modes into task definitions, contributor qualifications, calibration sets, quality controls, review workflows, and acceptance criteria.
Evaluation systems Established
Design evaluation logic, reviewer protocols, functionality tests, and claim-verification workflows that make decisions inspectable rather than merely plausible.
Human quality systems Established
Build calibration, second-level review, adjudication, and escalation for work that depends on expert judgment. Disagreement is a diagnostic signal, not automatically noise.
Data partnerships Current direction
Provider qualification, provenance review, pilot design, and acceptance criteria for specialized-data sourcing -- a trust and translation layer between model teams and data owners.
Proprietary data for AI Current direction
Help make internal knowledge and operational data usable as governed, testable inputs for retrieval, fine-tuning, evaluation, and agent workflows.
Agent environments & RL In development
Long-horizon tasks built from real or measured data, with hidden variation, explicit baselines, and reproducible grading.
The Problems I Want to Work On Next
These are research and operating questions, not a list of packaged services. I treat utility, provenance, rights, and measurement as separate claims that need separate evidence.
- What evidence should support a claim that a dataset is audited or fit for purpose?
- How should a model team specify data utility before acquisition rather than discovering its requirements after delivery?
- How should technical utility, provenance confidence, licensing scope, and maintainability be evaluated without collapsing them into one quality score?
- How can scarce domain expertise become reusable, interactive training signal through agent environments and reinforcement-learning tasks?
- Where should model-assisted quality review stop and calibrated human adjudication begin?
- How can organizations turn proprietary knowledge into governed, testable infrastructure for retrieval, fine-tuning, evaluation, and agent workflows?
- What open specifications and audit records would make private or paid data markets more transparent without requiring the underlying data to be public?
Discuss a Difficult Data or Evaluation Problem
I am most interested in research-data operations, specialized-data sourcing, marketplace quality, proprietary-data systems, human oversight, and agent evaluation. A useful first message names the objective, the available evidence, and the part of the system that is currently uncertain.
✦ Learning Voyage LLC is my independent practice and the publishing home for this work.