Being-H0.8: A Latent Tactile World-Action Model at Scale

Jul 28, 2026
BeingBeyond Team
Paper
Contents

Introduction

We introduce Being-H0.8, our latest general-purpose embodied foundation model. Being-H0.8 extends latent world-action modeling from visual prediction to tactile-aware interaction and, to our knowledge, is the first model to bring tactile pretraining to egocentric human video at this scale.

0:00 / 0:00

Visual world models can anticipate how scenes evolve, but vision alone cannot fully capture the contact dynamics that govern dexterous manipulation. Critical decisions arise at the contact interface: whether contact has been established, which hand region bears the load, and how the hand should adapt afterward. These cues are often subtle or occluded in RGB observations, yet they can determine the success of grasping, insertion, twisting, and handovers. Being-H0.8 therefore treats touch as an action-relevant part of the shared latent world state.

The central challenge is scale. Physical tactile sensors remain costly and are largely absent from existing robot and human data. Being-H0.8 instead adopts hand–surface interaction as a unified tactile representation, turning unlabeled human manipulation video into a scalable source of contact and proximity supervision. We introduce TactoHand, trained on geometry-supervised human–object interaction data, to generate dense pseudo-tactile annotations. Binary contact maps localize likely contact regions, while continuous proximity fields capture how close each hand vertex is to contact. When available, robot and gloved-human collections additionally provide measured contact or pressure, with validity masks specifying which tactile and action dimensions are observed in each sample.

Being-H0.8 visuo-tactile world-action model overview
Figure 1. Overview of Being-H0.8 across internet-scale visuo-tactile pretraining, task-aligned human mid-training, and robot post-training.

To connect this tactile supervision with action learning across embodiments, Being-H0.8 introduces TopoHand, a skeleton-based representation that replaces the first-generation Unified Action Space of Being-H0.5. TopoHand provides a shared state–action language for human hands, dexterous robot hands, and parallel grippers. Each hand is represented by a canonical wrist-centered frame, a compact morphology latent, and canonical articulation variables. Parallel grippers enter the same space through a morphology-normalized opening value derived from the distance between the thumb and its nearest fingertip. This design lets Being-H0.8 learn hand trajectories and robot controls in shared action slots, without embodiment-specific policy heads, and improves transfer from human pretraining to dexterous manipulation.

Real-World Contact-Rich Manipulation

These real-world demonstrations probe continuous force modulation, deformable-object handling, contact coordination, rapid adaptation, and precise tool use. Each task is shown from a single focused view so that its contact sequence remains easy to follow.

Hardware Platforms

Being-H0.8 is deployed on two complementary bimanual platforms. The first uses two ROKAE AR5 manipulators and supports interchangeable dexterous hands for cross-embodiment evaluation: the LinkerHand L25, with 16 actuated and 21 total degrees of freedom, and the tendon-driven DexHand 021, with 12 actuated and 22 total degrees of freedom. Each L25 fingertip carries a 12 × 6 piezoresistive pressure array, while the DexHand 021 integrates fingertip visuotactile sensors. The second platform combines two RealMan RM65 manipulators with Inspire grippers and Daimon DM-TAC W2M tactile sensors, providing a complementary setting for bimanual coordination and contact-aware control.

Squeeze Toothpaste

Deformable objectForce controlTouch feedback
Squeezing a deformable tube requires coordinated grasp geometry and continuously adjusted force as the tube empties.

Retrieve Hidden Items

Occluded retrievalIn-bag searchBimanual
Touch guides bimanual search and retrieval inside a soft bag when its contents are largely hidden from view.

Brush Calligraphy

Fine tool controlStroke precisionContact control
Writing with a soft brush demands precise tool orientation, stable contact, and finely controlled stroke pressure.

Pluck a Grape

Delicate forceContact precisionFragile object
The robot applies just enough force to separate one grape without crushing it or disturbing the surrounding bunch.

Wipe Away a Stain

Surface contactPressure control
Removing a stubborn stain requires sustained surface contact, broad coverage, and consistent pressure.

Adaptive Fruit Handling

Force adaptationVarying softness
The robot adjusts its grip across fruits of different firmness, securing each one without bruising softer pieces.

Grasp a Moving Cake

Dynamic sceneFast reactionDelicate force
The robot tracks a moving target, reacts quickly, and uses gentle force to lift the soft cake without deforming it.

Pick Potato Chips

Brittle objectFingertip precision
Thin, brittle chips require precise fingertip placement and enough force to lift without crushing.

Clean a Changing Spill

Human interventionAdaptive cleanup
The robot maintains contact and replans as a person adds fresh liquid midway through the cleanup.

Adaptive Shoe Polishing

Human interventionCurved contactTool use
The robot maintains controlled pressure over a curved surface and continues smoothly through human intervention.

Touch the World: Entering an Era of 500,000+ Hours

UniHand 3.0 provides the data foundation for Being-H0.8, drawing on 500,000+ hours of egocentric human video from public collections and a broad network of data partners. TactoHand recovers tactile information absent from conventional RGB video by converting observed hand–object interactions into dense contact and proximity supervision. We complement this corpus with standardized heterogeneous robot data, synthetic human–robot aligned demonstrations, and task-aligned multimodal human demonstrations. Together, these sources teach Being-H0.8 not only how the physical world evolves, but also how interaction unfolds at the point of contact.

Egocentric Human Data

To train Being-H0.8, we assemble 500,000+ hours of raw human video. To the best of our knowledge, this is the first human-centric corpus at this scale for which every sample is traceable to its source.

UniHand 3.0

Pre-training data atlas

From raw video to retained training data

UniHand 3.0500,000+raw hours

Filter by source

500,000+ hours

A large-scale, diverse collection spanning complementary data sources.

Sector angle · Raw duration
Inner area · Retained data
Figure 2. UniHand 3.0 brings together 500,000+ hours of traceable human-centric video from a broad mixture of open collections, partner-contributed data, and complementary self-built sources. Sector angle indicates raw duration, while the inner fill indicates data retained after curation.

The corpus draws on open egocentric collections, contributions from many data partners, and complementary self-built sources. This mixture expands coverage across activities, objects, environments, viewpoints, and recording conditions without relying on any single collection channel.

These sources vary substantially in format, annotation schema, camera configuration, and intended use. We unify them through a semi-automated, human-in-the-loop pipeline. Automated tools handle large-scale ingestion, metadata extraction, format conversion, and filtering; human reviewers audit each source and validate source-specific processing procedures. The result is a consistent training interface over otherwise heterogeneous data.

Across the Being-H series, human-video pretraining has progressed through three stages. Being-H0 established that human video contains transferable priors for robotic manipulation. Being-H0.5 and Being-H0.7 then expanded the scale and diversity of human interactions available for pretraining. Being-H0.8 marks a third stage, shifting the central question from how much video can be collected to which video provides reliable and useful embodied supervision. We address this challenge with a comprehensive pipeline for cleaning, deduplicating, filtering, standardizing, and structuring egocentric video at scale.

Data preparation and curation
From raw egocentric video to reliable embodied supervision
Video preparation, deduplication, and filtering pipelineA Sankey-style flow from an unprocessed egocentric video pool through validity checks, multiple deduplication stages, automated quality filters, and final manual screening.Unprocessed video poolUnprocessedvideo poolCorrupted video detectionCorrupted videodetectionNon-action filteringNon-actionfilteringFuzzy deduplicationFuzzydeduplicationExact deduplicationExactdeduplicationSemantic deduplicationSemanticdeduplicationUnstable camera motion filteringUnstablecamera motionfilteringLow-quality hand trajectory filteringLow-qualityhand trajectoryfilteringHeuristic filteringHeuristicfilteringManual screeningManualscreeningValid event detectionDeduplicationFiltering
Figure 3. The UniHand 3.0 preparation pipeline schematically transforms raw egocentric video into reliable embodied supervision through validity checks, complementary deduplication stages, automated quality filtering, and final manual screening.

The system goes beyond conventional video filtering. It converts raw egocentric footage into structured embodied supervision, including interaction-centric clips, temporally consistent hand trajectories, and rich 3D grounding signals. Nominal scale alone does not determine training value: relevance, quality, redundancy, and grounding checks substantially reshape the raw collection before training. The next bottleneck in embodied pretraining is therefore not data acquisition alone, but the reliable discovery, reconstruction, and validation of valuable experience within raw human video.

TactoHand: Scalable Tactile Labeling for Human Video

Egocentric video captures physical interaction at scale but does not directly record touch. TactoHand recovers this missing signal from temporal visual evidence. Given a tracked hand clip, it predicts two dense fields over the canonical MANO surface: a binary field that localizes contact and a continuous field that measures proximity to contact. Together, they describe both the contact event and its onset.

TactoHand replaces costly tactile annotation with geometric supervision. Across diverse hand–object interaction frames, registered hand and object meshes provide surface-aligned targets. Inter-surface distance and orientation determine contact, while distance also defines a graded proximity signal. A temporal model learns to infer both fields from video, using motion context to resolve interactions that remain ambiguous in a single frame.

UniHand 3.0 scales TactoHand to unlabeled egocentric video by reconstructing and tracking each hand, applying the model to the resulting clips, and mapping predictions back to the shared MANO topology. Unreliable reconstructions are masked rather than mislabeled as no contact. The result is dense, topology-aligned pseudo-tactile supervision for the universal tactile encoder.

TactoHand Data Samples

Where measured tactile signals are available, device-specific anatomical mappings project sparse contact and pressure observations onto the same canonical hand surface. Contact, proximity, and pressure therefore share one spatial interface without conflating their distinct meanings or validity.

Standardizing Heterogeneous Robot Data

Robot data provides a critical bridge between human-video pretraining and real-world deployment. Yet robot datasets differ widely across embodiments and collection pipelines, often using incompatible joint-axis and sign conventions, end-effector definitions, coordinate frames, camera setups, and state–action representations. Combining them without calibration introduces systematic inconsistencies into embodied pretraining.

To manage this heterogeneity, we separate each source into a reusable platform-level embodiment specification and dataset-level metadata. Standardization then becomes two tractable steps: mapping robot platforms to a shared embodiment interface, followed by dataset-specific calibration, synchronization, and validation.

Robot data standardization across embodiment modeling, dataset alignment, filtering, and augmentation
Figure 4. Heterogeneous robot data are standardized through platform-level embodiment modeling, dataset-level calibration and reprojection, and quality-controlled filtering and augmentation.

For each supported platform, a canonical URDF defines arm kinematics, bimanual geometry, joint conventions, and wrist or end-effector frames. Calibration parameters recovered from robot geometry, recorded metadata, and visual observations map raw states and actions into this shared specification. Actions are then expressed as wrist poses in the camera coordinate frame.

Each aligned dataset is checked for data integrity, cross-stream synchronization, kinematic consistency, and language–video agreement. Failed samples are filtered or routed through dedicated reprocessing, while high-quality episodes are augmented only with geometry-consistent transformations that preserve alignment among images, states, actions, and language.

Synthetic Human–Robot Aligned Data

Large-scale robot foundation models require broad and diverse manipulation data, yet collecting real robot demonstrations is expensive, difficult to scale, and constrained by the task distribution of each collection setup. Human egocentric video offers rich manipulation knowledge, but raw recordings also contain irrelevant motion, occlusion, unstable interactions, and environmental noise that complicate hand reconstruction and robot action transfer.

We therefore use generated human manipulation video as an intermediate representation between human demonstrations and robot learning. Unlike unconstrained real-world recordings, generated videos can be controlled through textual and visual conditions to produce cleaner, more task-focused hand–object interactions. This improves the reliability of 3D hand reconstruction, motion retargeting, and robot control while enabling scalable variation across objects, scenes, and tasks.

Our synthesis pipeline uses two complementary strategies. Text-to-video provides broad coverage of object–action combinations, while image-to-video conditions generation on visual context to preserve scene structure and interaction consistency. An automatic evaluator checks visual quality, instruction alignment, and physical consistency, removing samples with visual degradation, semantic mismatch, scale errors, or implausible motion.

Scalable synthesis of paired human egocentric and robot manipulation data
Figure 5. Scalable human-to-robot paired data synthesis combines diverse egocentric video generation with hand-motion reconstruction, embodiment-aware retargeting, and depth-aware robot rendering.

The retained human videos are then converted into executable robot demonstrations. The pipeline reconstructs 3D hand poses and motion trajectories; removes human body regions through hand and arm segmentation, occlusion-aware inpainting, and object restoration; and transfers the recovered motion to multiple robot embodiments through finger retargeting, palm and fingertip alignment, and arm inverse kinematics. Robot-specific constraints are enforced while preserving the original manipulation intent.

Finally, depth-aware compositing combines robot renderings with the reconstructed scene to produce synchronized RGB images, depth observations, and action trajectories. This automated path from controllable human-video generation to multi-embodiment robot data provides a scalable, lower-cost source of paired demonstrations for foundation-model pretraining.

Task-Aligned Human Multimodal Demonstrations

We collect real-world bimanual human demonstrations as task-aligned training data. Performed directly by humans, these demonstrations match the downstream task distribution and provide relevant manipulation semantics, motion strategies, and natural contact behavior without the cost and constraints of large-scale robot data collection.

This dataset bridges broad multimodal pretraining and robot-specific adaptation. Task alignment focuses the model on relevant object interactions and temporal structure, while human demonstrations provide dexterous, contact-rich behaviors that are difficult to capture with robots alone. The resulting supervision strengthens task understanding, hand–object interaction representations, temporal prediction, and tactile sensitivity before adaptation to robot action spaces.

Toward a Latent Tactile World-Action Model

Being-H0.8 extends the latent World–Action Model of Being-H0.7 from visual prediction to tactile-aware interaction learning. Four components make this possible: a latent tactile world-action formulation that learns from future visual and tactile evidence; a slow–fast action expert that combines persistent world-action context with frequent state and tactile updates; a universal tactile encoder that unifies heterogeneous tactile signals across datasets and embodiments; and TopoHand, which provides a shared action representation for human hands, dexterous robot hands, and grippers.

Being-H0.8 prior-posterior tactile world-action model pipeline
Figure 6. Being-H0.8 aligns deployable visuo-tactile prior embeddings with a future-informed posterior under a shared Mixture-of-Transformers backbone and branch-aware attention mask.

Prior–Posterior Formulation

Being-H0.8 retains the prior–posterior formulation of Being-H0.7. The deployable prior combines the current instruction and visual observations with learned latent queries, while the training-only posterior fills the corresponding latent positions with future visual and tactile evidence.

Because the two branches share the same token layout, their corresponding latent states can be aligned directly. The future-informed posterior supervises the prior, encouraging it to anticipate action-relevant latent states from the currently available context.

Rather than reconstructing future observations at the pixel level, Being-H0.8 captures their task-relevant consequences in latent space and uses the predicted states to condition action generation. At inference time, the posterior and all future observations are removed, leaving a deployable latent world model.

Slow-Fast Action Expert

Standard action experts reduce inference cost by predicting a complete action chunk from the current context and executing it largely open loop. This is poorly suited to contact-rich manipulation, where upcoming actions must often change as new proprioceptive and tactile feedback becomes available.

Being-H0.8 slow-fast action expert with state and tactile feedback anchors
Figure 7. The slow–fast action expert preserves full-horizon prediction while refreshing near-term action segments with the latest state and tactile tokens.

At the beginning of each action chunk, the slow stream computes and caches visual–language and latent world context while predicting the complete horizon. Before each shorter segment is executed, the fast stream reuses that context, incorporates the latest proprioceptive state and tactile feedback, and regenerates the upcoming segment.

The two streams share action-expert parameters and targets. A blockwise causal attention pattern restricts each fast segment to feedback available by its execution anchor, and only the next segment is committed. This preserves full-horizon planning with higher-frequency reactive control.

Universal Tactile Encoder

Tactile observations vary widely across datasets and embodiments, from global contact labels and sparse fingertip sensors to dense per-vertex contact, proximity, and pressure measurements. We organize this heterogeneity as a coarse-to-fine tactile pyramid over a canonical hand topology, with each level describing the same hand at a different spatial granularity. During broad pretraining, one valid level is sampled for each example; embodiment-specific training and deployment use the level that best matches the available annotation or sensor layout.

Universal tactile encoder from multi-granularity hand signals to fixed tactile tokens
Figure 8. The universal tactile encoder maps heterogeneous contact, proximity, and pressure signals to a shared fixed-token interface.

The selected tactile locations and measurements are embedded as a variable-length token sequence. A query-based Perceiver resamples that sequence into a fixed number of tactile tokens with a shared feature dimension. When tactile sensing is unavailable, a learned missing-touch representation preserves the same interface, allowing the Being-H0.8 posterior and slow–fast action expert to consume touch independently of the original sensor density or layout.

TopoHand: A Unified Kinematic State–Action Space

Human and robot hands differ in geometry, joint configuration, and native control space, but share common skeletal semantics. TopoHand turns this structure into a unified state–action representation whose slots retain the same anatomical meaning across embodiments. Unlike MANO-specific policy outputs, these shared slots do not change meaning between human-video pretraining and robot adaptation.

TopoHand unified action space and skeleton shape variational autoencoder
Figure 9. TopoHand expresses human hands, dexterous robot hands, and grippers in one canonical state–action space while factoring skeleton morphology into a compact latent code.

TopoHand separates morphology from articulation. A shape VAE compresses the canonical rest-hand skeleton into a ten-dimensional morphology code, while wrist motion and 20 canonical joint angles describe the current state and future action. Embodiment-specific adapters convert native human or robot motion into this shared space and map predicted actions back to native commands.

For each hand, the state comprises wrist translation and rotation, the morphology code, and canonical joint angles. The action comprises wrist translation and rotation updates followed by absolute target joint angles. Bimanual samples concatenate the representations for both hands.

The same interface supports parallel grippers. A morphology-normalized opening signal is derived from the shortest distance between the thumb tip and the remaining fingertips. Adding one opening coordinate per side gives human hands, dexterous robot hands, and grippers a single shared policy interface.

Together, the four components define a consistent information path. Future visual and tactile observations supervise a deployable latent representation; the slow–fast action expert reuses that representation while incorporating current state and touch to refine actions in real time. The universal tactile encoder standardizes heterogeneous physical observations, while TopoHand standardizes states and actions across embodiments.

Quality Analysis

Egocentric Human Data

The nominal scale of an embodied dataset often overstates its effective diversity. This is especially severe in simulation, where repeated environments, task templates, initial states, and control policies generate visually and behaviorally similar trajectories. Real-robot datasets are likewise constrained by a limited range of embodiments, workspaces, and collection protocols. Egocentric human data spans a much broader distribution of activities and environments, but raw footage still contains substantial semantic and temporal redundancy, with highly variable relevance and quality. Our curation pipeline removes redundant episodes and segments, bringing UniHand 3.0 closer to the diversity profile of general-purpose multimodal video while preserving the manipulation-centric structure required for embodied learning.

Diversity and deduplication
UniHand 3.0
LLaVA-Video-178K
Sim Robot Data
Real Robot Data
Raw Human Data
(a) Episodes remaining0%25%50%75%100%.001.03.07.124.165.205.26UniHand 3.0: 99.99% at ε=0.00095UniHand 3.0: 86.00% at ε=0.03UniHand 3.0: 50.45% at ε=0.07UniHand 3.0: 18.87% at ε=0.124UniHand 3.0: 8.65% at ε=0.165UniHand 3.0: 4.30% at ε=0.205UniHand 3.0: 1.61% at ε=0.26LLaVA-Video-178K: 99.43% at ε=0.00095LLaVA-Video-178K: 97.61% at ε=0.03LLaVA-Video-178K: 86.09% at ε=0.07LLaVA-Video-178K: 54.80% at ε=0.124LLaVA-Video-178K: 32.90% at ε=0.165LLaVA-Video-178K: 18.87% at ε=0.205LLaVA-Video-178K: 8.27% at ε=0.26Sim Robot Data: 99.91% at ε=0.00095Sim Robot Data: 1.06% at ε=0.03Sim Robot Data: 0.50% at ε=0.07Sim Robot Data: 0.15% at ε=0.124Sim Robot Data: 0.09% at ε=0.165Sim Robot Data: 0.09% at ε=0.205Sim Robot Data: 0.06% at ε=0.26Real Robot Data: 96.74% at ε=0.00095Real Robot Data: 47.96% at ε=0.03Real Robot Data: 12.02% at ε=0.07Real Robot Data: 2.79% at ε=0.124Real Robot Data: 1.05% at ε=0.165Real Robot Data: 0.43% at ε=0.205Real Robot Data: 0.16% at ε=0.26Raw Human Data: 99.61% at ε=0.00095Raw Human Data: 60.06% at ε=0.03Raw Human Data: 28.16% at ε=0.07Raw Human Data: 8.07% at ε=0.124Raw Human Data: 3.52% at ε=0.165Raw Human Data: 1.67% at ε=0.205Raw Human Data: 0.51% at ε=0.26Deduplication threshold (ε)
(b) Videos with duplicates0%25%50%75%100%.001.03.07.124.165.205.26UniHand 3.0: 0.01% at ε=0.00095UniHand 3.0: 16.59% at ε=0.03UniHand 3.0: 58.87% at ε=0.07UniHand 3.0: 91.22% at ε=0.124UniHand 3.0: 97.68% at ε=0.165UniHand 3.0: 99.33% at ε=0.205UniHand 3.0: 99.90% at ε=0.26LLaVA-Video-178K: 0.57% at ε=0.00095LLaVA-Video-178K: 2.80% at ε=0.03LLaVA-Video-178K: 17.62% at ε=0.07LLaVA-Video-178K: 59.17% at ε=0.124LLaVA-Video-178K: 83.31% at ε=0.165LLaVA-Video-178K: 94.18% at ε=0.205LLaVA-Video-178K: 98.79% at ε=0.26Sim Robot Data: 0.09% at ε=0.00095Sim Robot Data: 99.68% at ε=0.03Sim Robot Data: 99.91% at ε=0.07Sim Robot Data: 100.00% at ε=0.124Sim Robot Data: 100.00% at ε=0.165Sim Robot Data: 100.00% at ε=0.205Sim Robot Data: 100.00% at ε=0.26Real Robot Data: 3.36% at ε=0.00095Real Robot Data: 58.98% at ε=0.03Real Robot Data: 92.46% at ε=0.07Real Robot Data: 99.09% at ε=0.124Real Robot Data: 99.78% at ε=0.165Real Robot Data: 99.94% at ε=0.205Real Robot Data: 99.99% at ε=0.26Raw Human Data: 0.41% at ε=0.00095Raw Human Data: 44.70% at ε=0.03Raw Human Data: 81.38% at ε=0.07Raw Human Data: 96.82% at ε=0.124Raw Human Data: 99.08% at ε=0.165Raw Human Data: 99.71% at ε=0.205Raw Human Data: 99.98% at ε=0.26Deduplication threshold (ε)
(c) Episode-level maximum similarity0%10%20%30%40%50%0.60.70.80.91.0UniHand 3.0: 0.00% at similarity 0.61UniHand 3.0: 0.01% at similarity 0.63UniHand 3.0: 0.00% at similarity 0.65UniHand 3.0: 0.01% at similarity 0.67UniHand 3.0: 0.02% at similarity 0.69UniHand 3.0: 0.02% at similarity 0.71UniHand 3.0: 0.04% at similarity 0.73UniHand 3.0: 0.11% at similarity 0.75UniHand 3.0: 0.19% at similarity 0.77UniHand 3.0: 0.39% at similarity 0.79UniHand 3.0: 0.62% at similarity 0.81UniHand 3.0: 1.28% at similarity 0.83UniHand 3.0: 2.38% at similarity 0.85UniHand 3.0: 4.96% at similarity 0.87UniHand 3.0: 8.29% at similarity 0.89UniHand 3.0: 13.89% at similarity 0.91UniHand 3.0: 19.06% at similarity 0.93UniHand 3.0: 22.04% at similarity 0.95UniHand 3.0: 18.58% at similarity 0.97UniHand 3.0: 8.11% at similarity 0.99LLaVA-Video-178K: 0.01% at similarity 0.61LLaVA-Video-178K: 0.02% at similarity 0.63LLaVA-Video-178K: 0.05% at similarity 0.65LLaVA-Video-178K: 0.09% at similarity 0.67LLaVA-Video-178K: 0.16% at similarity 0.69LLaVA-Video-178K: 0.33% at similarity 0.71LLaVA-Video-178K: 0.55% at similarity 0.73LLaVA-Video-178K: 0.97% at similarity 0.75LLaVA-Video-178K: 1.62% at similarity 0.77LLaVA-Video-178K: 2.89% at similarity 0.79LLaVA-Video-178K: 4.83% at similarity 0.81LLaVA-Video-178K: 7.31% at similarity 0.83LLaVA-Video-178K: 10.77% at similarity 0.85LLaVA-Video-178K: 14.36% at similarity 0.87LLaVA-Video-178K: 16.12% at similarity 0.89LLaVA-Video-178K: 15.91% at similarity 0.91LLaVA-Video-178K: 11.62% at similarity 0.93LLaVA-Video-178K: 7.10% at similarity 0.95LLaVA-Video-178K: 3.88% at similarity 0.97LLaVA-Video-178K: 1.22% at similarity 0.99Real Robot Data: 0.00% at similarity 0.61Real Robot Data: 0.00% at similarity 0.63Real Robot Data: 0.00% at similarity 0.65Real Robot Data: 0.00% at similarity 0.67Real Robot Data: 0.00% at similarity 0.69Real Robot Data: 0.00% at similarity 0.71Real Robot Data: 0.00% at similarity 0.73Real Robot Data: 0.01% at similarity 0.75Real Robot Data: 0.02% at similarity 0.77Real Robot Data: 0.04% at similarity 0.79Real Robot Data: 0.06% at similarity 0.81Real Robot Data: 0.12% at similarity 0.83Real Robot Data: 0.25% at similarity 0.85Real Robot Data: 0.57% at similarity 0.87Real Robot Data: 1.27% at similarity 0.89Real Robot Data: 2.76% at similarity 0.91Real Robot Data: 6.26% at similarity 0.93Real Robot Data: 14.59% at similarity 0.95Real Robot Data: 41.63% at similarity 0.97Real Robot Data: 32.43% at similarity 0.99Raw Human Data: 0.00% at similarity 0.61Raw Human Data: 0.00% at similarity 0.63Raw Human Data: 0.00% at similarity 0.65Raw Human Data: 0.00% at similarity 0.67Raw Human Data: 0.00% at similarity 0.69Raw Human Data: 0.01% at similarity 0.71Raw Human Data: 0.01% at similarity 0.73Raw Human Data: 0.04% at similarity 0.75Raw Human Data: 0.11% at similarity 0.77Raw Human Data: 0.16% at similarity 0.79Raw Human Data: 0.24% at similarity 0.81Raw Human Data: 0.43% at similarity 0.83Raw Human Data: 0.92% at similarity 0.85Raw Human Data: 1.64% at similarity 0.87Raw Human Data: 2.90% at similarity 0.89Raw Human Data: 6.86% at similarity 0.91Raw Human Data: 12.85% at similarity 0.93Raw Human Data: 18.68% at similarity 0.95Raw Human Data: 22.89% at similarity 0.97Raw Human Data: 32.26% at similarity 0.99Maximum within-cluster cosine similarity
Figure 10. Semantic redundancy across human and robot video sources. Panels (a) and (b) show the episodes retained and the videos containing at least one duplicate episode as the deduplication threshold increases. Panel (c) compares episode-level maximum within-cluster similarity. After curation, UniHand 3.0 moves away from the redundancy profile of raw egocentric and robot data toward the broader distribution of general-purpose multimodal video.

Diversity alone is insufficient when the underlying hand trajectories are unreliable. We therefore develop a scalable pipeline to reconstruct, stabilize, standardize, and validate hand motion from unconstrained egocentric video. Existing reconstruction systems often suffer from tracking drift, implausible geometry, motion discontinuities, and silent failures under occlusion or rapid camera movement. Designed for large-scale in-the-wild data, our pipeline substantially reduces severe tracking errors, hand-shape distortions, and unusable trajectories. Its automated quality-control stage further detects over 92% of erroneous episodes before pretraining.

Human motion reconstruction
5,000 sampled episodes
(a) Severe tracking failures0%25%50%75%100%67%HaWoR57%Dyn-HamR39%Ours
(b) Hand distortion / abnormal shape0%25%50%75%100%28%HaWoR17%Dyn-HamR11%Ours
(c) Failed after the full pipeline0%25%50%75%100%73%HaWoR61%Dyn-HamR11%Ours
(d) Erroneous episodes filtered0%25%50%75%100%0%Without QC92%With QC
Figure 11. Human evaluation of hand-motion reconstruction quality over 5,000 randomly sampled episodes. The Being-H0.8 reconstruction and quality-control pipeline reduces severe tracking failures, hand distortion, and full-pipeline failures, while filtering 92% of erroneous episodes before pretraining.

Robot Data

Unlike egocentric human video, where redundancy is the primary concern, robot-data quality is also shaped by heterogeneity. Existing collections span different embodiments, camera systems, control interfaces, and annotation pipelines, making nominal scale a poor proxy for pretraining value. Even visually similar collections may differ substantially in task coverage, state–action reliability, and geometric consistency. We therefore evaluate robot data along four complementary axes: state–action integrity, episode-level geometric consistency, task diversity, and visual diversity. Rather than collapsing these signals into a single score, we compare normalized quality profiles across sources. Some sources offer broader task or visual coverage, while others are cleaner but much narrower. Our curation and alignment pipeline reconciles these differences, transforming heterogeneous robot sources into a more diverse, geometrically consistent, and pretraining-ready corpus.

Robot data quality
Single-source datasetOURSBubble area · Relative data scale
(a)State-action integrity vs. geometric consistency
50%60%70%80%90%100%0%25%50%75%100%RoboCOIN: State-action integrity 74.6, Geometric consistency 58.5RoboCOINRoboMIND: State-action integrity 61.3, Geometric consistency 96.7RoboMINDInterndata-A1: State-action integrity 51.7, Geometric consistency 93.1Interndata-A1AgiBotWorld-Beta: State-action integrity 82.3, Geometric consistency 79.8AgiBotWorld-BetaGalaxea Open-World: State-action integrity 98.6, Geometric consistency 71.1Galaxea Open-WorldOURS: State-action integrity 100.0, Geometric consistency 100.0OURSState-action integrityGeometric consistency
(b)Task diversity vs. visual diversity
0510152004590135180RoboCOIN: Task diversity 33.4, Visual diversity 16.1RoboCOINRoboMIND: Task diversity 13.6, Visual diversity 7.6RoboMINDInterndata-A1: Task diversity 18.9, Visual diversity 15.3Interndata-A1AgiBotWorld-Beta: Task diversity 25.4, Visual diversity 13.1AgiBotWorld-BetaGalaxea Open-World: Task diversity 77.6, Visual diversity 17.1Galaxea Open-WorldOURS: Task diversity 166.6, Visual diversity 18.5OURSTask diversityVisual diversity
(c)Visual coverage in feature space
DINOv2 feature-space coverage across representative robot data sources
Figure 12. Robot-data quality across representative open-source datasets. Panels (a) and (b) compare state–action integrity with geometric consistency and task diversity with visual diversity, respectively; bubble area indicates relative source scale. Panel (c) visualizes DINOv2 feature-space coverage, highlighting the core sources while grouping the remainder as Other. Paired counts denote source-dominant and occupied visual clusters, using p(d|k) ≥ 0.8 for dominance.

Acknowledgements

We thank the following companies for serving as core data providers and supplying high-quality data for model training:

帕西尼感知科技、无问智科、星际硅途、星忆智能、枢途科技、智域基石、EgoScale、宽凳科技、博登智能、睿尔曼智能、数灵科技Digients Tech、五维数据

Citation

@misc{beingbeyond2026beingh08,
  title={{Being-H0.8}: A Latent Tactile World-Action Model at Scale},
  author={{BeingBeyond Team}},
  year={2026},
  howpublished={BeingBeyond Technical Report},
  url={https://research.beingbeyond.com/being-h08}
}