Home
Own the outcome
Evaluation protocol
If you cannot tell whether data is useful, you are selling garbage and labs are training on garbage. Video pretraining at scale is not the same as language pretraining. We own the protocol first — uplift numbers only when a real pilot produces them.
Principles
- Usefulness = task success under deployment-like conditions.
- Every Pro pack names ≥1 eval task and easy / medium / hard splits.
- Failure and recovery demos are first-class when tagged.
- Researcher-led pilots beat procurement-first claims.
Steps
- Define the task and success criteria with the buyer.
- Hold out eval clips; label difficulty; include failures when available.
- Measure baseline success on the buyer's policy or an open checkpoint.
- Fine-tune / train with the WA pack (document layers L0–L3 used).
- Re-measure; report Δ success and remaining failure modes.
- Feed failures into the next capture queue.
Public stance
Pilot results: TBD until signed evidence. We will not publish fabricated scaling plots. See the data pack template for how eval hooks are recorded per delivery.