# HumanRelay — Training-Ready Data Pipelines for Physical AI > HumanRelay provides training-ready data pipelines for physical AI: first-person 4K video of real people doing real work, captured and annotated end-to-end for VLA (vision-language-action) models, robots, and embodied AI. More than 3,000 hours of 4K video are captured daily, with a path to 7,000+ hours per day by the end of September, delivered by 1,500+ full-time employees — not gig workers. HumanRelay turns real-world activity into training-ready data for physical AI models. It specializes in egocentric data collection (first-person, head-mounted video) — the real-world, first-person data that VLA models and robots learn from. Annotation and evaluation are run with our sister company IndiVillage, which has delivered 500M+ annotations to global AI teams. Contact: hello@humanrelay.com. ## Egocentric data collection Egocentric data collection. First-person 4K video of real people doing real work — hands, tools, homes, kitchens, warehouses. We run managed capture centers where trained collectors — on staff, not crowdsourced — record first-person video of real tasks: cooking, cleaning, folding, assembly, picking, repair. The volume model builders need, annotated at scale by a workforce you can name. - Over 3,000 hours of 4K egocentric video captured daily, growing fast. - Head-mounted 4K rigs with synced audio + IMU. - Scripted task coverage or naturalistic capture. - Scripted task programs against your taxonomy, or naturalistic full-day capture. - Consent-first, full provenance metadata. - Capture and annotation under one roof — no vendor hand-off. - We go business to business, in person, to onboard and train teams to capture how real work actually gets done — on iPhones, approved devices, or customer-specific rigs. ### Capture specification - Rig: Head-mounted 4K @ 30/60 fps. - Streams: Video · audio · IMU. - Delivery: Raw · curated · annotation-ready. - Fidelity: 4K + LiDAR depth, true metric geometry. Crisp 4K imagery, much of it paired with LiDAR-derived depth and real-world measurements — the exact 3D geometry, and the real distances between hands and objects, a robot needs to act, not merely perceive. ### Why egocentric data matters VLA (vision-language-action) models and robots learn from first-person, egocentric data of how real work actually gets done. That data can't be scraped from the internet — it has to be captured in the physical world, where the work happens. Capturing it is slow, manual, unglamorous work — which is why, we believe, so little of it exists today. ## Scale & traction - ~1,000 businesses in our network — contracted or being contracted. - +3,000 hours of egocentric video captured every day. - 7,000+ hours/day — capture trajectory by the end of September. - Capture and annotation are delivered as one managed, training-ready data pipeline. ## Annotation (with IndiVillage) - 500M+ annotations delivered for global AI teams. - Video, sensor, and 3D/pose labeling. - RLHF preference ranking & rubric evals. - Dual-pass QA at 99%+ accuracy. - Enriched for how robot policies learn: hand-object interactions, action sequences, and dense manipulation signal. ## Workforce - 1,500+ full-time employees — not gig workers. - 94% college-educated · 1 in 3 with engineering or other technical degrees. - 99%+ QA accuracy · 98% client retention. - 15 offices across seven Indian states; global headquarters in London. - The seven operating states collectively represent nearly 700 million people and around 7 million registered businesses. ## How we compare (public estimates) - Meta (Ego4D): egocentric research video — a one-time academic collection. ~3,670 hrs. - Build AI (Egocentric-1M): crowd-captured glasses video at volume — but lower-resolution, raw, and shipped with no annotation. ~1M hrs (reported). - Physical Intelligence (humanoid / VLA lab): proprietary teleoperation on their own robots. ~10,000+ hrs (π0, reported). - HumanRelay (on the ground): 4K egocentric video from contracted, consented businesses — captured and annotated end-to-end by our in-house team. +3,000 hrs/day and growing 50–100%/wk. Raw hours aren't the moat — annotated, train-ready hours are. At our run-rate (+3,000 hrs/day) we add Ego4D's entire ~3,670-hour corpus in under two days, fully annotated and captured across real businesses. ## The human stack - [Capture](https://humanrelay.com/#capture): +3,000 hours of 4K egocentric video every day. Head-mounted 4K at 30/60 fps with synced audio and IMU, covering real manipulation tasks. Consent-first with provenance metadata and PII scrubbing. - [Annotate](https://humanrelay.com/#annotate): Annotation and evaluation proven at scale via IndiVillage — 500M+ annotations delivered. - [Judge](https://humanrelay.com/#judge): An API for human judgment. Your agents call it when they shouldn't guess; sub-minute resolution, tiered by complexity. - [Operate](https://humanrelay.com/#operate): Teleoperation for deployed robot fleets. A trained operator is on task in seconds; the full trace returns via API as an annotated demonstration. ## Key pages - [The stack](https://humanrelay.com/#the-stack): One partner. The whole human stack. - [Flywheel](https://humanrelay.com/#flywheel): Every intervention is a labeled demonstration. - [Workforce](https://humanrelay.com/#workforce): Managed teams. Not gig workers. 1,500+ full-time professionals. - [Footprint](https://humanrelay.com/#footprint): 15 offices across seven Indian states, with global reach rooted in India. - [Egocentric data collection](https://humanrelay.com/egocentric-data-collection): First-person 4K data capture for robots, VLA models, and embodied AI. - [Pricing](https://humanrelay.com/#pricing): Four products. Simple meters. - [FAQ](https://humanrelay.com/#faq): Questions, answered straight. ## Contact - Email: hello@humanrelay.com - Start a pilot: Judge pilots return results in 48 hours; annotation pilots start from as few as 10 hours of work; capture pilots deliver a sample batch against your taxonomy; teleop pilots begin with a scoped coverage window.