World Archive

The global data infrastructure for physical AI

Robots have no internetto learn from.So we're building it.

When a robotics team has a data problem, we're the answer to it. What to capture. Getting it captured. Labelling it. Auditing what you already have. Proving it made the policy better. All of it, from one partner — because we help robots learn better, not just record more footage.

  1. 01StrategyWhat to capture — and when not to
  2. 02CaptureStereo RGB-D, wrist, tactile
  3. 03Annotate & VerifyOur footage or your existing data
  4. 04EvaluateMeasure the uplift

Stereo RGB-D capture. Dual wrist cameras. Tactile gloves. 21-point hand pose. Object tracks. Action segments. Sub-frame sync. 6-DoF calibration. Human-verified QA. Consent-first. LeRobot · Hugging Face · MCAP. India · Uzbekistan · Türkiye · Romania. ~2,000 hrs/week capacity

01Why policies plateau

Your model has neverwatched anyoneactually work.

Language models had the whole web. Physical AI has almost nothing — no internet-scale record of how skilled work gets done. The open corpora that exist (Ego4D, Open X-Embodiment, DROID) are a sliver, and almost none of it is multi-modal capture of real jobs. It can't be scraped, because it was never online.

01

The tasks you need aren't in any dataset

Bench demos and lab teleop cover a narrow band of clean behaviour. The long tail — a heat gun, a sewing bench, a service bay, a nursing round — is where deployment actually happens, and where public data runs out entirely.

Your policy generalises to the demo and fails on the job.

02

You can't tell whether the data helped

Most data purchases end the same way: buy hours, fine-tune, hope. No held-out task, no baseline, no measurement. When the policy doesn't improve, nothing tells you whether it was the data, the split, or the recipe.

You spend the budget and learn nothing you can act on.

02One partner, every data need

Whatever the dataproblem is,it's ours.

Most vendors sell you hours of footage and leave the rest to you. We take the whole problem — including the data you already have, and including telling you when you don't need more of it.

  • 01Data strategy

    You don't know what to capture

    We map which task families actually move your policy, in what order, before anyone signs a purchase order.

  • 02Field capture

    The data you need doesn't exist

    Multi-modal rigs and crews on consented sites across four countries, scoped to your tasks.

  • 03Annotation at volume

    You have footage, not labels

    Hands, objects, actions and contact — on our capture or on the footage you already own.

  • 04Quality audit

    You don't trust the data you have

    We assess sync, calibration, label accuracy and coverage on an existing corpus and tell you what is trainable.

  • 05Evaluation

    You don't know if it worked

    Held-out tasks and a real baseline, so improvement is measured instead of assumed.

  • 06Sim-ready delivery

    You need it in simulation

    Metric depth and per-session camera models, so packs import without a cleanup project.

We help robots learn better.

That is the whole point. Footage is just the part you can see.

03What we do

Six ways in.One partner.

Teams normally stitch together a strategy consultant, a capture vendor, an annotation shop and an eval harness — then carry the integration risk themselves. Every one of these is ours, and you can start at whichever one you need.

  1. 01

    Data strategy

    Available now

    We start by telling you what you actually need.

    Most teams don't need more data yet — they need to know which task families move their policy and in what order. We map that with you before any purchase order, and we'll say so when the answer is 'not yet'.

    • Which task families to prioritise
    • Modality trade-offs for your embodiment
    • A sequenced roadmap, not a quote
    • Team upskilling on capture and sync
  2. 02

    Field capture

    Available now

    The data that doesn't exist yet.

    We deploy purpose-built rigs on consented sites across four countries — stereo RGB-D on the head, cameras on both wrists, tactile gloves where contact matters. Sub-frame synced and calibrated every session, because geometry you can't trust isn't worth training on.

    • Stereo RGB-D head, dual wrist RGB, IMU
    • Tactile gloves for contact-rich tasks
    • Factory floors and household environments
    • Custom task families on request
  3. 03

    Annotation & QA

    Available now

    On our footage, or on yours.

    Models draft the layers; our team verifies every one. Blur, lost tracks, wrong joints and heavy occlusion are rejected before anything reaches you. We run the same pipeline on datasets you already own — you don't have to have bought the capture from us.

    • 21-point hand pose, both hands, per frame
    • Object boxes with persistent track IDs
    • Verb–noun action segments with timestamps
    • Applied to your existing corpus too
  4. 04

    Quality audit

    Available now

    Find out what your corpus is really worth.

    Bring us a dataset — ours, a vendor's or your own — and we report on what is actually trainable: sync integrity, calibration, label accuracy, coverage gaps and the share that should be rejected outright.

    • Sync and calibration integrity checks
    • Label accuracy sampling against ground truth
    • Coverage and difficulty gap analysis
    • A reject list, with reasons
  5. 05

    Evaluation

    In development

    Proof the data worked.

    Every pack can ship with a held-out task, difficulty splits and a baseline score, so improvement is measured rather than assumed. We're building this on OpenVLA as a public reference, and we can run evaluations on the same real sites the data came from.

    • Held-out eval tasks per pack
    • Easy / medium / hard splits
    • Baseline, then post-training re-measure
    • On-site evaluation at our capture facilities
  6. 06

    Simulation

    On the roadmap

    Scale past what cameras can record.

    Because we capture metric depth and log camera models every session, your packs already import into Isaac- and Omniverse-style pipelines without cleanup. Digital twins of the sites we capture, and synthetic generation across them, are what we build next.

    • Sim-ready geometry today
    • Digital twins of captured sites
    • Synthetic generation at volume
    • real2sim / sim2real validation
04How your robot gets better

A dataset is one round.Improvement is a loop.

Buying hours once and hoping is not a strategy. Capture what matters, structure it, measure what changed, then go get what's still missing. We run four of those five stages so your team only has to do the part it's good at.

  1. 01

    Capture

    We handle it

    We scope the task families with you, then record them properly — multi-modal, synced, on consented sites.

  2. 02

    Structure

    We handle it

    Sync, calibration, hand pose, object tracks, action segments, contact. Verified by people before it ships.

  3. 03

    Train

    You train

    Your policy, your stack. Packs arrive ready for LeRobot, Hugging Face and MCAP — no ingestion project.

  4. 04

    Measure

    We handle it

    Score a baseline, train on the pack, re-measure on held-out tasks. A number you can take to a review.

  5. 05

    Target the gaps

    We handle it

    Whatever still fails becomes the brief for the next capture. Each round is aimed better than the last.

The second round is always sharper than the first, because by then we know exactly what your policy is failing at.

05See it before you commission it

Real footage.Real labels.

No mock-ups on this page. Everything below is pulled from live capture programmes and shipped delivery layers.

Every modality, on real jobs

Stereo pairs, wrist views, tactile sessions and synced dual-view composites — the stacks teams ask to see before committing to volume.

Drag to seewhat ships.

One second of a real episode, three ways. Raw capture on the left, the delivery layer on the right — actual exports, not a render.

  1. 01Raw captureSource stream, untouched
  2. 02Hand pose21-point landmarks, both hands, per frame
  3. 03Object tracksBoxes with persistent track IDs
Compare against
RawHand pose

21-point landmarks, L/R, per frame

06What you receive

A pack yourpipeline can readon day one.

Every delivery ships the same contract: layered data, the metadata to trust it, and the files to load it. No bespoke parsing, no guessing what a column means.

Layers in every pack

  1. L0Synced streamsTime-aligned video from every camera on the rig, plus IMU and tactile where captured.
  2. L1Calibration & timingIntrinsics, extrinsics and per-frame timestamps, so multi-camera geometry actually resolves.
  3. L2Hands & objects21-point hand pose per frame, object boxes with persistent track IDs, overlay previews for review.
  4. L3Actions & contactVerb–noun segments with start/end times, grasp and release events, contact states.

Formats

  • LeRobot
  • Hugging Face Datasets
  • MCAP
  • MP4
  • JSONL
  • RLDS / OXE on request

Example pack

tree
pack/  DATACARD.md  videos/          head, wrist_l, wrist_r  depth/           metric maps + point clouds  annotations/    hand_pose.jsonl    objects.json    segments.jsonl    contact.jsonl  metadata/    calibration.json    frame_timestamps.csv    consent.json  qa/report.csv  manifest.sha256
07Choosing a capture stack

Why we capturethe expensive way.

Not every programme needs every modality, and we'll tell you when it doesn't. But the differences are physical, not marketing — here is what each choice costs you downstream.

  1. 01Measured geometry, not estimated

    Stereo over monocular

    One lens means depth and scale are inferred, not measured. Two lenses with a known baseline give metric depth — which is what grasp, contact and collision actually need, and what transfers into sim.

  2. 02Clean motion at speed

    Global shutter over rolling

    Rolling shutter reads the sensor line by line, so fast hand and tool motion skews. That warp pollutes pose, contact timing and stereo matching — quietly, in ways you find during training.

  3. 03Streams that mean the same instant

    Sub-frame sync over 'close enough'

    At 30–60 fps a frame is 16–33 ms, and a grasp changes state inside that window. Clap, QR and PTP discipline keep streams agreeing tighter than one frame, so contact labels land where the contact was.

  4. 04Calibration that holds all day

    Industrial rigs over phones

    Phone-and-cage rigs are heavy, fatigue operators and loosen mid-session. Loose mounts drift extrinsics silently, and multi-camera geometry is gone before anyone notices.

Monocular still has a place for high-volume density programmes — we'll scope that where it's the right call.

08Coverage

If people do it,we can capture it.

We run business development across thousands of workplaces, which means the long tail is reachable at a price that makes sense. If your task isn't listed, it's usually a scoping call, not a no.

Manufacturing

  • Steel & metals
  • Forging & casting
  • Fasteners
  • Light manufacturing
  • Packaging & kitting
  • Textile & sewing

Skilled trades

  • Welding
  • Carpentry & joinery
  • Painting & finishing
  • Automotive body work
  • Repair & maintenance

Retail & services

  • Grocery & kirana
  • Pharmacy
  • Sweets & food prep
  • Telecom
  • Automobile service bays

Craft & niche

  • Pottery
  • Cane weaving
  • Sculpture
  • Heritage craft
  • Leather & edge finishing
  • Mehndi & tattoo

Sports & recreation

  • Shuttlecock assembly
  • Cricket bats
  • Football stitching
  • Trophies
  • Packaging & QC

Care & clinical

  • Nursing workflows
  • Equipment preparation
  • Veterinary
  • Consent-first, PII-cleared

Household

  • Kitchen tasks
  • Laundry & folding
  • Cleaning & mopping
  • Tidying & sorting
Not listed?We scope custom task families on request.Tell us the task
09Quality & rights

Built to clearyour review.

The thing that stalls a data deal is rarely the footage — it's whether legal, security and ML all sign off. That gets designed in.

How quality is enforced

  1. 01

    Human reject-QA

    Models draft, people verify. Bad frames are culled before delivery, and the QA report ships with the pack.

  2. 02

    Calibration discipline

    Intrinsics and extrinsics per session, re-calibrated hourly and after any remount.

  3. 03

    Provenance per layer

    Every layer tagged human, model, derived or mixed. You always know what was measured and what was inferred.

  4. 04

    Reproducible packs

    SHA256 manifests and a DATACARD, so a pack you got last quarter is the pack you can still audit.

Rights & governance

Consent-first capture
Operators and site owners consent before any commercial capture. Consent metadata travels with the data.
Clear usage rights
Training, evaluation and redistribution defined per contract. Exclusive and non-exclusive tiers available.
PII reduction
Faces, screens, badges and plates flagged, blurred or rejected during verification.
Controlled delivery
Checksummed packs, access-controlled transfer, delivery logs on every handover.
10How we work

Start small.Prove it.Then scale.

No six-figure commitment before you've seen whether the data works. Engagements are built to give you a real answer early.

  1. 01Week 1

    Scope

    A founder-led session on the behaviour you need, the environment, the modalities and how you'll judge success. You leave with a capture protocol and a difficulty ladder.

  2. 02Weeks 2–4

    Pilot

    A small scoped pack — typically 100–500 hrs/week — delivered against the contract, with the QA report and eval hooks. Enough to train on and judge properly.

  3. 03After delivery

    Review

    We go through label quality, coverage and what your policy did with it. Anything that missed gets corrected in the protocol, not argued about.

  4. 04Ongoing

    Scale

    Ramp toward available capacity (~2,000 hrs/week) with the protocol locked, on an exclusive or non-exclusive programme.

100–500hrs/week
Typical pilot volume
~0hrs/week
Available capacity
0countries
India · Uzbekistan · Türkiye · Romania
0+rigs
Purpose-built field kits
11Where this goes

Every robot that workslearned from somewhere.

We intend to be that somewhere — the record of how physical work actually gets done, and the standard teams measure their policies against. Capture first, because it's the part nobody else wanted to do properly.

12Talk to the founders

Tell us the task.We'll tell youwhat it takes.

Send the behaviour you need a robot to perform, the environment it happens in, and how you'll know it worked. You'll get a straight answer on feasibility, modalities and timeline — from the people who run the capture, not a sales desk.

What to expect

  • Founder-led scoping, no SDR round-trip
  • Advisory before any purchase order
  • Exclusive and non-exclusive programmes
  • Custom task families on request