Este sitio utiliza cookies. Si continúa navegando en este sitio, acepta nuestro uso de cookies. Lea nuestra política de privacidad>

Búsqueda
  • NLCHRIC: A Solid Data Foundation for High-Performance Humanoid Robots

    NLCHRIC: A Solid Data Foundation for High-Performance Humanoid Robots

"How could it be so clear and cool?
For fresh water comes from the source."
These words, written centuries ago by the poet Zhu Xi, now take on new meaning as engineers search for the source of life for humanoid robots.

Inside a fully enclosed training facility in Shanghai, a humanoid robot repeatedly performs the same actions: identifying a target, grasping it, and precisely placing it. This is the daily scene at the core site of Humanoid Robot (Shanghai) Co., Ltd., also known as the National and Local Co-built Humanoid Robotics Innovation Center (NLCHRIC).

These seemingly monotonous actions generate valuable data for the humanoid robots, supporting their application in real-world scenarios. In high-end manufacturing, robots work in automotive manufacturing plants to tackle tasks like parts assembly and inspections. In urban services, they are being piloted for street cleaning and waste collection. In terms of emergency response, robots will be able to deliver supplies, perform rescue demolition, and enter fire grounds to conduct reconnaissance.

As the first national-level innovation center for the humanoid robotics industry, jointly established by the Ministry of Industry and Information Technology (MIIT) and the Shanghai Municipal Government, the NLCHRIC handles multiple national key projects. It also pursues the strategic mission of building "1+N" training facilities and a data foundation, aiming to successfully transition humanoid robots from the lab to large-scale, real-world implementation.

Data Challenges: "Growing Pains" of Humanoid Robots

The full lifecycle of robot training incorporates a comprehensive process, from data acquisition and governance to annotation and circulation. Building high-quality datasets is especially critical—but this step involves numerous challenges.

First, massive data volume places extreme demands on storage. Each training session generates petabytes of multi-modal data, including video, point clouds, and force signals, posing severe challenges to storage capacity, performance, and total cost of ownership (TCO). Storing such massive amounts of data economically and efficiently has therefore emerged as a primary challenge.

Second, data acquisition is costly, as high-quality data requires reliable transfer and long-term storage. Real-world data acquisition involves significant equipment investment and lengthy acquisition cycles. Any data loss or corruption can result in substantial losses. This makes a stable and reliable transfer network essential, while robust storage mechanisms are equally critical for long-term data retention.

Third, the complexity and diversity of data formats make data governance challenging. Data from vision, tactile sensing, force sensing, and pose estimation can significantly vary in format. Without an efficient data management system, cleansing, annotating, and integrating data become difficult, thus limiting data utilization for model training. Meanwhile, a limited range of data acquisition methods, paired with missing samples and high acquisition costs, make it difficult to scale datasets.

Fourth, difficulties remain regarding cross-region data sharing. Operational data from real-world scenarios is subject to complex data security and ownership constraints, raising obstacles to efficient data reuse across organizations and projects, and limiting data value.

These numerous data-related challenges are currently barriers to the development of humanoid robots.

To address these challenges, the NLCHRIC has adopted a phased approach. First, it is building a reliable storage foundation for its training facilities, ensuring sufficient storage capacity and efficient data management. Second, the NLCHRIC is transitioning to an AI data lake and plans to build the foundation model data acquisition and training platform, which will enable cross-region data aggregation, intelligent scheduling, and efficient data circulation, thus making data truly dynamic.

Solid Data Foundation: A "Super Granary" for Training Facilities

To accelerate the transition of humanoid robots from single-point capability validation to sustained real-world operation and iterative improvement, the NLCHRIC has built an embodied-intelligence technology framework featuring a "5+4+1" structure, meaning five large models, four data middle platforms, and one dataset.

Guided by scenario-specific requirements and starting with an evaluation of application value and technical feasibility, the framework covers the entire workflow—from scenario evaluation, task definition, data acquisition, data validation and governance, and dataset accumulation to model training, comprehensive evaluation, scenario-specific deployment, operational data feedback, and iterative optimization. This creates a self-reinforcing technical closed loop in which data, models, evaluation, and applications drive one another.

And putting this framework into practice starts with the data layer. The NLCHRIC has adopted Huawei OceanStor Pacific scale-out storage system to build a unified storage foundation for training facilities. The core capabilities of this system span three dimensions.

First, it delivers high-concurrency throughput to support rapid loading of large-scale datasets and real-time data ingestion, reducing data preparation time before model iteration. Second, it provides massive scalability for the secure, long-term retention of historical data such as raw video and point clouds, supporting repeated training, algorithm traceability, and compliance retention. Third, with a unified management view, data can be stored in tiers based on use cases and centrally orchestrated, eliminating fragmented storage. So far, a 6-PB resource pool has been established, providing ample storage capacity.

The value of this storage system goes far beyond simply storing data. It serves as the hub of the entire data closed loop, enabling multi-protocol interworking across the data lake. A single copy of data supports multi-task concurrent access across protocols such as S3, NFS, HDFS, and POSIX. It also adaptively processes large and small I/Os across hybrid workloads, allowing data to move seamlessly from acquisition and desensitization to upload.

Building on this foundation, hot and cold data are automatically tiered, with a global orchestration view. Data can move among training facilities on demand and be flexibly scheduled across on-premises and cloud environments, transforming data management from static storage to dynamic data flow. At the same time, the full-stack lightweight cloud solution provides unified management of AI servers, storage, networks, virtualization, and containers. Capabilities such as data engineering, model engineering, and agent development can all be deployed as needed, enabling efficient collaboration between storage systems and training platforms.

Powered by this storage infrastructure, data seamlessly flows into the labeling and governance pipeline. Through the synergy of five model types, the integration of four platforms, and the continuous enrichment of the Baihu dataset 2.0, the NLCHRIC is steadily building an embodied AI closed-loop technology system—anchored by scenarios, built on data, powered by models, gated by evaluation, and propelled by application feedback. This system provides comprehensive support for data aggregation, model R&D, capability evaluation, and scenario deployment across the city's humanoid robotics industry ecosystem.

Taking this progress a step further, the NLCHRIC and Huawei have jointly launched the first showcase for embodied intelligence training in China. Building on the full-link closed-loop framework at the demo center, the two parties will align and continuously feed back real-time and high-fidelity simulation data in both directions. Together, they are constructing a comprehensive data flywheel—spanning heterogeneous acquisition terminals, parallel pipelines for simulation data generation, industry data standards, and an integrated data governance platform—expanding the training facility from technical depth to operational scale.

Channeling Data into Lakes: Navigating Toward Industry-Wide Transformation

While the "super granary" addresses the challenges of data storage and management, NLCHRIC's vision goes further: building a full-stack AI data lake that bridges training facilities with the real world. Now that the technical closed loop is up and running, the next step is extending it into a business closed loop for actual commercial value.

According to the NLCHRIC's plan, the foundation model data acquisition and training platform serves as the central hub where data from edge facilities converges to form an all-region data lake. The scenario-specific middle platform then deploys these capabilities into industrial and service environments, completing the flywheel with real-world feedback.

This vision is no longer out of reach. Currently, the NLCHRIC has signed agreements with six parties, including the Shanghai Municipal Fire and Rescue Bureau and State Grid Shanghai Municipal Electric Power Company, covering six core sectors such as emergency response and energy and power. These partnerships are accelerating the deployment of humanoid robots for routine operations in industrial environments. Meanwhile, the Shanghai Humanoid Robotics Pilot-Scale Service Platform is now fully operational in Pudong district. With an annual production capacity of 2,000 complete robot units, it provides solid support for large-scale deployment.

These tangible achievements are turning humanoid robots from an aspiration into reality, making them productive partners and capable assistants.

And the starting point of it all lies behind every grasp and every step. From robust data infrastructure to a dynamic AI data lake transfer, and ultimately to multi-industry application implementation, it is this end-to-end pipeline—integrating data acquisition, storage, governance, training, and deployment—that gives cold metallic skeletons a sense of perception and empowers silicon-based neural networks to truly learn and adapt.

Arriba