Introduction
We introduce Being-H0.8, our latest general-purpose embodied foundation model. Being-H0.8 extends latent world-action modeling from visual prediction to tactile-aware interaction and, to our knowledge, is the first model to bring tactile pretraining to egocentric human video at this scale.
Visual world models can anticipate how scenes evolve, but vision alone cannot fully capture the contact dynamics that govern dexterous manipulation. Critical decisions arise at the contact interface: whether contact has been established, which hand region bears the load, and how the hand should adapt afterward. These cues are often subtle or occluded in RGB observations, yet they can determine the success of grasping, insertion, twisting, and handovers. Being-H0.8 therefore treats touch as an action-relevant part of the shared latent world state.
The central challenge is scale. Physical tactile sensors remain costly and are largely absent from existing robot and human data. Being-H0.8 instead adopts hand–surface interaction as a unified tactile representation, turning unlabeled human manipulation video into a scalable source of contact and proximity supervision. We introduce TactoHand, trained on geometry-supervised human–object interaction data, to generate dense pseudo-tactile annotations. Binary contact maps localize likely contact regions, while continuous proximity fields capture how close each hand vertex is to contact. When available, robot and gloved-human collections additionally provide measured contact or pressure, with validity masks specifying which tactile and action dimensions are observed in each sample.

To connect this tactile supervision with action learning across embodiments, Being-H0.8 introduces TopoHand, a skeleton-based representation that replaces the first-generation Unified Action Space of Being-H0.5. TopoHand provides a shared state–action language for human hands, dexterous robot hands, and parallel grippers. Each hand is represented by a canonical wrist-centered frame, a compact morphology latent, and canonical articulation variables. Parallel grippers enter the same space through a morphology-normalized opening value derived from the distance between the thumb and its nearest fingertip. This design lets Being-H0.8 learn hand trajectories and robot controls in shared action slots, without embodiment-specific policy heads, and improves transfer from human pretraining to dexterous manipulation.
Real-World Contact-Rich Manipulation
These real-world demonstrations probe continuous force modulation, deformable-object handling, contact coordination, rapid adaptation, and precise tool use. Each task is shown from a single focused view so that its contact sequence remains easy to follow.
Hardware Platforms
Being-H0.8 is deployed on two complementary bimanual platforms. The first uses two ROKAE AR5 manipulators and supports interchangeable dexterous hands for cross-embodiment evaluation: the LinkerHand L25, with 16 actuated and 21 total degrees of freedom, and the tendon-driven DexHand 021, with 12 actuated and 22 total degrees of freedom. Each L25 fingertip carries a 12 × 6 piezoresistive pressure array, while the DexHand 021 integrates fingertip visuotactile sensors. The second platform combines two RealMan RM65 manipulators with Inspire grippers and Daimon DM-TAC W2M tactile sensors, providing a complementary setting for bimanual coordination and contact-aware control.
Squeeze Toothpaste
Retrieve Hidden Items
Brush Calligraphy
Pluck a Grape
Wipe Away a Stain
Adaptive Fruit Handling
Grasp a Moving Cake
Pick Potato Chips
Clean a Changing Spill
Adaptive Shoe Polishing
Touch the World: Entering an Era of 500,000+ Hours
UniHand 3.0 provides the data foundation for Being-H0.8, drawing on 500,000+ hours of egocentric human video from public collections and a broad network of data partners. TactoHand recovers tactile information absent from conventional RGB video by converting observed hand–object interactions into dense contact and proximity supervision. We complement this corpus with standardized heterogeneous robot data, synthetic human–robot aligned demonstrations, and task-aligned multimodal human demonstrations. Together, these sources teach Being-H0.8 not only how the physical world evolves, but also how interaction unfolds at the point of contact.
Egocentric Human Data
To train Being-H0.8, we assemble 500,000+ hours of raw human video. To the best of our knowledge, this is the first human-centric corpus at this scale for which every sample is traceable to its source.
UniHand 3.0
Pre-training data atlas
From raw video to retained training data
Filter by source
500,000+ hours
A large-scale, diverse collection spanning complementary data sources.
The corpus draws on open egocentric collections, contributions from many data partners, and complementary self-built sources. This mixture expands coverage across activities, objects, environments, viewpoints, and recording conditions without relying on any single collection channel.
These sources vary substantially in format, annotation schema, camera configuration, and intended use. We unify them through a semi-automated, human-in-the-loop pipeline. Automated tools handle large-scale ingestion, metadata extraction, format conversion, and filtering; human reviewers audit each source and validate source-specific processing procedures. The result is a consistent training interface over otherwise heterogeneous data.
Across the Being-H series, human-video pretraining has progressed through three stages. Being-H0 established that human video contains transferable priors for robotic manipulation. Being-H0.5 and Being-H0.7 then expanded the scale and diversity of human interactions available for pretraining. Being-H0.8 marks a third stage, shifting the central question from how much video can be collected to which video provides reliable and useful embodied supervision. We address this challenge with a comprehensive pipeline for cleaning, deduplicating, filtering, standardizing, and structuring egocentric video at scale.
The system goes beyond conventional video filtering. It converts raw egocentric footage into structured embodied supervision, including interaction-centric clips, temporally consistent hand trajectories, and rich 3D grounding signals. Nominal scale alone does not determine training value: relevance, quality, redundancy, and grounding checks substantially reshape the raw collection before training. The next bottleneck in embodied pretraining is therefore not data acquisition alone, but the reliable discovery, reconstruction, and validation of valuable experience within raw human video.
TactoHand: Scalable Tactile Labeling for Human Video
Egocentric video captures physical interaction at scale but does not directly record touch. TactoHand recovers this missing signal from temporal visual evidence. Given a tracked hand clip, it predicts two dense fields over the canonical MANO surface: a binary field that localizes contact and a continuous field that measures proximity to contact. Together, they describe both the contact event and its onset.
TactoHand replaces costly tactile annotation with geometric supervision. Across diverse hand–object interaction frames, registered hand and object meshes provide surface-aligned targets. Inter-surface distance and orientation determine contact, while distance also defines a graded proximity signal. A temporal model learns to infer both fields from video, using motion context to resolve interactions that remain ambiguous in a single frame.
UniHand 3.0 scales TactoHand to unlabeled egocentric video by reconstructing and tracking each hand, applying the model to the resulting clips, and mapping predictions back to the shared MANO topology. Unreliable reconstructions are masked rather than mislabeled as no contact. The result is dense, topology-aligned pseudo-tactile supervision for the universal tactile encoder.
TactoHand Data Samples
Where measured tactile signals are available, device-specific anatomical mappings project sparse contact and pressure observations onto the same canonical hand surface. Contact, proximity, and pressure therefore share one spatial interface without conflating their distinct meanings or validity.
Standardizing Heterogeneous Robot Data
Robot data provides a critical bridge between human-video pretraining and real-world deployment. Yet robot datasets differ widely across embodiments and collection pipelines, often using incompatible joint-axis and sign conventions, end-effector definitions, coordinate frames, camera setups, and state–action representations. Combining them without calibration introduces systematic inconsistencies into embodied pretraining.
To manage this heterogeneity, we separate each source into a reusable platform-level embodiment specification and dataset-level metadata. Standardization then becomes two tractable steps: mapping robot platforms to a shared embodiment interface, followed by dataset-specific calibration, synchronization, and validation.

For each supported platform, a canonical URDF defines arm kinematics, bimanual geometry, joint conventions, and wrist or end-effector frames. Calibration parameters recovered from robot geometry, recorded metadata, and visual observations map raw states and actions into this shared specification. Actions are then expressed as wrist poses in the camera coordinate frame.
Each aligned dataset is checked for data integrity, cross-stream synchronization, kinematic consistency, and language–video agreement. Failed samples are filtered or routed through dedicated reprocessing, while high-quality episodes are augmented only with geometry-consistent transformations that preserve alignment among images, states, actions, and language.
Synthetic Human–Robot Aligned Data
Large-scale robot foundation models require broad and diverse manipulation data, yet collecting real robot demonstrations is expensive, difficult to scale, and constrained by the task distribution of each collection setup. Human egocentric video offers rich manipulation knowledge, but raw recordings also contain irrelevant motion, occlusion, unstable interactions, and environmental noise that complicate hand reconstruction and robot action transfer.
We therefore use generated human manipulation video as an intermediate representation between human demonstrations and robot learning. Unlike unconstrained real-world recordings, generated videos can be controlled through textual and visual conditions to produce cleaner, more task-focused hand–object interactions. This improves the reliability of 3D hand reconstruction, motion retargeting, and robot control while enabling scalable variation across objects, scenes, and tasks.
Our synthesis pipeline uses two complementary strategies. Text-to-video provides broad coverage of object–action combinations, while image-to-video conditions generation on visual context to preserve scene structure and interaction consistency. An automatic evaluator checks visual quality, instruction alignment, and physical consistency, removing samples with visual degradation, semantic mismatch, scale errors, or implausible motion.

The retained human videos are then converted into executable robot demonstrations. The pipeline reconstructs 3D hand poses and motion trajectories; removes human body regions through hand and arm segmentation, occlusion-aware inpainting, and object restoration; and transfers the recovered motion to multiple robot embodiments through finger retargeting, palm and fingertip alignment, and arm inverse kinematics. Robot-specific constraints are enforced while preserving the original manipulation intent.
Finally, depth-aware compositing combines robot renderings with the reconstructed scene to produce synchronized RGB images, depth observations, and action trajectories. This automated path from controllable human-video generation to multi-embodiment robot data provides a scalable, lower-cost source of paired demonstrations for foundation-model pretraining.
Task-Aligned Human Multimodal Demonstrations
We collect real-world bimanual human demonstrations as task-aligned training data. Performed directly by humans, these demonstrations match the downstream task distribution and provide relevant manipulation semantics, motion strategies, and natural contact behavior without the cost and constraints of large-scale robot data collection.
This dataset bridges broad multimodal pretraining and robot-specific adaptation. Task alignment focuses the model on relevant object interactions and temporal structure, while human demonstrations provide dexterous, contact-rich behaviors that are difficult to capture with robots alone. The resulting supervision strengthens task understanding, hand–object interaction representations, temporal prediction, and tactile sensitivity before adaptation to robot action spaces.
Toward a Latent Tactile World-Action Model
Being-H0.8 extends the latent World–Action Model of Being-H0.7 from visual prediction to tactile-aware interaction learning. Four components make this possible: a latent tactile world-action formulation that learns from future visual and tactile evidence; a slow–fast action expert that combines persistent world-action context with frequent state and tactile updates; a universal tactile encoder that unifies heterogeneous tactile signals across datasets and embodiments; and TopoHand, which provides a shared action representation for human hands, dexterous robot hands, and grippers.

Prior–Posterior Formulation
Being-H0.8 retains the prior–posterior formulation of Being-H0.7. The deployable prior combines the current instruction and visual observations with learned latent queries, while the training-only posterior fills the corresponding latent positions with future visual and tactile evidence.
Because the two branches share the same token layout, their corresponding latent states can be aligned directly. The future-informed posterior supervises the prior, encouraging it to anticipate action-relevant latent states from the currently available context.
Rather than reconstructing future observations at the pixel level, Being-H0.8 captures their task-relevant consequences in latent space and uses the predicted states to condition action generation. At inference time, the posterior and all future observations are removed, leaving a deployable latent world model.
Slow-Fast Action Expert
Standard action experts reduce inference cost by predicting a complete action chunk from the current context and executing it largely open loop. This is poorly suited to contact-rich manipulation, where upcoming actions must often change as new proprioceptive and tactile feedback becomes available.

At the beginning of each action chunk, the slow stream computes and caches visual–language and latent world context while predicting the complete horizon. Before each shorter segment is executed, the fast stream reuses that context, incorporates the latest proprioceptive state and tactile feedback, and regenerates the upcoming segment.
The two streams share action-expert parameters and targets. A blockwise causal attention pattern restricts each fast segment to feedback available by its execution anchor, and only the next segment is committed. This preserves full-horizon planning with higher-frequency reactive control.
Universal Tactile Encoder
Tactile observations vary widely across datasets and embodiments, from global contact labels and sparse fingertip sensors to dense per-vertex contact, proximity, and pressure measurements. We organize this heterogeneity as a coarse-to-fine tactile pyramid over a canonical hand topology, with each level describing the same hand at a different spatial granularity. During broad pretraining, one valid level is sampled for each example; embodiment-specific training and deployment use the level that best matches the available annotation or sensor layout.

The selected tactile locations and measurements are embedded as a variable-length token sequence. A query-based Perceiver resamples that sequence into a fixed number of tactile tokens with a shared feature dimension. When tactile sensing is unavailable, a learned missing-touch representation preserves the same interface, allowing the Being-H0.8 posterior and slow–fast action expert to consume touch independently of the original sensor density or layout.
TopoHand: A Unified Kinematic State–Action Space
Human and robot hands differ in geometry, joint configuration, and native control space, but share common skeletal semantics. TopoHand turns this structure into a unified state–action representation whose slots retain the same anatomical meaning across embodiments. Unlike MANO-specific policy outputs, these shared slots do not change meaning between human-video pretraining and robot adaptation.

TopoHand separates morphology from articulation. A shape VAE compresses the canonical rest-hand skeleton into a ten-dimensional morphology code, while wrist motion and 20 canonical joint angles describe the current state and future action. Embodiment-specific adapters convert native human or robot motion into this shared space and map predicted actions back to native commands.
For each hand, the state comprises wrist translation and rotation, the morphology code, and canonical joint angles. The action comprises wrist translation and rotation updates followed by absolute target joint angles. Bimanual samples concatenate the representations for both hands.
The same interface supports parallel grippers. A morphology-normalized opening signal is derived from the shortest distance between the thumb tip and the remaining fingertips. Adding one opening coordinate per side gives human hands, dexterous robot hands, and grippers a single shared policy interface.
Together, the four components define a consistent information path. Future visual and tactile observations supervise a deployable latent representation; the slow–fast action expert reuses that representation while incorporating current state and touch to refine actions in real time. The universal tactile encoder standardizes heterogeneous physical observations, while TopoHand standardizes states and actions across embodiments.
Quality Analysis
Egocentric Human Data
The nominal scale of an embodied dataset often overstates its effective diversity. This is especially severe in simulation, where repeated environments, task templates, initial states, and control policies generate visually and behaviorally similar trajectories. Real-robot datasets are likewise constrained by a limited range of embodiments, workspaces, and collection protocols. Egocentric human data spans a much broader distribution of activities and environments, but raw footage still contains substantial semantic and temporal redundancy, with highly variable relevance and quality. Our curation pipeline removes redundant episodes and segments, bringing UniHand 3.0 closer to the diversity profile of general-purpose multimodal video while preserving the manipulation-centric structure required for embodied learning.
Diversity alone is insufficient when the underlying hand trajectories are unreliable. We therefore develop a scalable pipeline to reconstruct, stabilize, standardize, and validate hand motion from unconstrained egocentric video. Existing reconstruction systems often suffer from tracking drift, implausible geometry, motion discontinuities, and silent failures under occlusion or rapid camera movement. Designed for large-scale in-the-wild data, our pipeline substantially reduces severe tracking errors, hand-shape distortions, and unusable trajectories. Its automated quality-control stage further detects over 92% of erroneous episodes before pretraining.
Robot Data
Unlike egocentric human video, where redundancy is the primary concern, robot-data quality is also shaped by heterogeneity. Existing collections span different embodiments, camera systems, control interfaces, and annotation pipelines, making nominal scale a poor proxy for pretraining value. Even visually similar collections may differ substantially in task coverage, state–action reliability, and geometric consistency. We therefore evaluate robot data along four complementary axes: state–action integrity, episode-level geometric consistency, task diversity, and visual diversity. Rather than collapsing these signals into a single score, we compare normalized quality profiles across sources. Some sources offer broader task or visual coverage, while others are cleaner but much narrower. Our curation and alignment pipeline reconciles these differences, transforming heterogeneous robot sources into a more diverse, geometrically consistent, and pretraining-ready corpus.

Acknowledgements
We thank the following companies for serving as core data providers and supplying high-quality data for model training:
Citation
@misc{beingbeyond2026beingh08,
title={{Being-H0.8}: A Latent Tactile World-Action Model at Scale},
author={{BeingBeyond Team}},
year={2026},
howpublished={BeingBeyond Technical Report},
url={https://research.beingbeyond.com/being-h08}
}
