RLHF
Human preference data when a checker is not enough. Pairwise judgments, ranked outputs, and reward signals from people — for the parts of the work that still need taste, context, and disagreement.
Environments, verifiers, and human data for training and evaluating models. Each line is built so correctness can be checked — or, where it cannot yet, so we can move the boundary.
Human preference data when a checker is not enough. Pairwise judgments, ranked outputs, and reward signals from people — for the parts of the work that still need taste, context, and disagreement.
Pre-existing datasets for post-training and fine-tuning, ready to use rather than commissioned from scratch. Demonstrations, traces, and labeled examples that bootstrap a model before RL — SFT, continued training, and the data a lab needs on day one.
Task worlds a model can act in, with a checker attached. Success is computed, not argued after the fact. We design the environment, the verifier, and the reliability number together.
Reinforcement learning for open-ended work where a structured, scalar signal can still be verified. For domains on the boundary — writing, investigation, and other tasks where pass/fail is not a single string.
Scorecards that turn an output into a grade. Criteria, weights, and checkers for when the right answer is a standard of quality, not a unique token sequence.
Reinforcement learning from verifiable rewards. Train and evaluate where a program can say yes or no: code, tools, math, and other domains with a ground-truth checker.