What experts score
Typical item shapes:
Prompts may be chat turns, tool-using agent traces, voice transcripts, or environment step logs. Stay inside the instructions for that project.
Decision standards
Use this order when comparing outputs:- Safety and truth: Wrong clinical/financial/legal facts, unsafe advice, or leaked private data lose immediately.
- Task success: Did the response actually solve the user or environment goal?
- Domain procedure: Did it follow the process a specialist would expect (confirmations, escalation, required checks)?
- Clarity and usefulness: Prefer concise, actionable answers over fluff, only after 1 to 3 are equal.
- Style: Tone and formatting are last; never let polish beat correctness.
How to write rationales
Keep rationales short and falsifiable:- Name the decisive defect or strength (e.g. “missed red-flag symptom”, “correctly required confirmation before place order”).
- Quote or paraphrase the specific span of the response that drove the decision.
- Avoid vague praise (“more helpful”, “sounds better”) unless you can say in what domain sense.
Consistency tips
- Read the project rubric every session: criteria differ by domain pack.
- If unsure, use the project’s “needs review / escalate” path instead of guessing.
- Do not invent knowledge outside the prompt and allowed references.
- Grade the response as shown, not what you wish the model had said after an imagined follow-up.
What “good” looks like for RLHF
Submissions are useful when:- Preferences agree with other experts on clear cases (high inter-rater agreement).
- Hard cases include a crisp rationale another specialist can audit.
- Safety failures are never ranked above safe-but-incomplete answers.
- Experts skip or escalate out-of-domain items instead of improvising.
Handoff into Evaluation
Approved RLHF work becomes preference and review signal your org uses with Evaluation:- Preference pairs and rankings for training and offline comparison
- Safety flags for gating and audit
- Rationales that support rubrics and human-review workflows
Related
- Reviewing agent runs: the core audit and review work type
- Domain writing for RL environments: when experts author the tasks agents are graded on
- Policies: no external LLMs on task content
- Accept assigned work: assignments and queues

