Este sitio utiliza cookies. Si continúa navegando en este sitio, acepta nuestro uso de cookies. Lea nuestra política de privacidad>

Búsqueda
  • The Data Bedrock Beneath the Furrows: Securing Food Autonomy Through Bio-breeding Innovation

    The Data Bedrock Beneath the Furrows: Securing Food Autonomy Through Bio-breeding Innovation

There is no doubt that, in terms of food self-sufficiency, China has long secured its ability to feed its people. A few months ago, a Canadian digital media company released an authoritative global ranking on food self-sufficiency, and China—home to 1.4 billion people—was already ranked third in the world.

But what we often see is only the outcome on the table. Looking across the entire agricultural value chain, the upstream breeding process that produces these results still faces significant challenges. In 2025, China's No. 1 Central Document once again emphasized the continued advancement of industrialized bio-breeding. This marks the fifth consecutive year that bio-breeding has been highlighted at the highest policy level—underscoring both its strategic importance and the urgency of innovation in the industry.

Agriculture sustains a nation, and seeds are its foundation. Seeds are the foundational technology of agriculture, the cornerstone of modern farming, and the original source of national food security. They are directly tied to people's livelihoods and the broader national economy.

To truly secure the rice bowl of China, the country must also cultivate its own superior seeds.

The good news is that this transformation is already underway. As a major national agricultural research institution, Yazhouwan National Laboratory is dedicated to seed science and breeding innovation. In collaboration with Huawei, it is actively leveraging AI and data-driven technologies to improve breeding efficiency and outcomes—accelerating the transformation toward intelligent and sustainable agriculture.

The Breeding Revolution Driven by AI for All—yet Hitting a Data Wall

At an industry event, experts from Yazhouwan National Laboratory reflected on the evolution of global plant breeding while sharing their AI practices in breeding. Like many traditional industries, breeding has long relied heavily on experience. Despite multiple stages of technological advancement, it has fundamentally remained a probability-driven discipline:

The first generation—domestication—relied on selecting superior seeds directly from nature.

The second generation—hybrid breeding—introduced strong randomness, requiring years of field trials, waiting, and observation in the hope of achieving favorable outcomes.

The third generation—molecular breeding—brought molecular markers into play, enabling pre-screening. However, at its core, it still only increased the probability of success in hybrid breeding.

The fourth generation—genomics-assisted breeding—evolved molecular markers into genetic testing and mapping, further improving efficiency and success rates.

It becomes clear that any approach unable to truly predict breeding outcomes in advance remains essentially constrained by probability.

This has created a time-based barrier. Countries that started earlier, relying on decades of trial-and-error and probability accumulation, have built advantages that are extremely difficult to catch up with in the short term. (From hybrid combinations to trait selection, breeding cycles often span years or even decades.) When China's breeding technologies were still at a generational disadvantage—and the probability of success lagged behind that of leading countries—the gap became even harder to close.

So what is the way forward?

If there were a method to break free from probability-driven breeding—if outcomes could be predicted before field trials even begin—then breeding would no longer be a game of incremental improvement, but a true leapfrog opportunity.

This is precisely the opportunity brought by fifth-generation breeding: intelligent variety design.

China's rapid progress in AI for All is driving efficiency upgrades—and in many cases, systemic transformation—across fields. AI is not merely a revolutionary tool, but a tool for revolution. In the field of breeding, the emerging paradigm of AI for Science enables AI models to rapidly analyze the relationships between genotype and phenotype, predict key traits such as yield and stress resistance, and dramatically shorten R&D cycles. This shift helps China's seed industry move beyond probability-driven limitations and accelerate toward technological catch-up and leadership.

However, as Yazhouwan National Laboratory embraced AI and actively moved into intelligent variety design, it quickly ran into a serious data challenge.

It is no secret that AI has shifted from the model-centric stage to a data-driven one. Data quality largely determines model performance. Many nations have elevated scientific data to the status of a national strategic asset in their AI strategies. Yet in China's agricultural sector—especially in the field of breeding—the data available today is still far from sufficient to support AI-driven breeding at scale. On the one hand, in terms of volume and distribution, agricultural data in China is highly scattered across different regions and difficult to share. On the other hand, in terms of quality, there is a lack of unified standards, with inconsistent data formats and quality metrics.

Against this backdrop, building data infrastructure that unifies data specifications and standards and lays a solid data foundation for AI applications has become a top priority for Yazhouwan National Laboratory.

Aggregation, Sharing, and Flow of High-Quality Datasets: a Data Symphony for Smart Breeding

In response to these deep-seated data challenges, Yazhouwan National Laboratory strategically partnered with Huawei to build an AI data lake foundation tailored for next-generation bio-breeding technologies.

To solve a problem effectively, one must first understand it deeply.

When it comes to building a data foundation, the first step is to establish a shared definition of what constitutes good data, that is, high-quality datasets. In collaboration between Yazhouwan National Laboratory and Huawei, the objectives of data development were structured into five progressive levels: starting from compliance with FAIR principles (Findable, Accessible, Interoperable, and Reusable), advancing toward support for general-purpose models, and then enabling enhanced inference capabilities, followed by building datasets specifically for science, and ultimately realizing plug-and-play unified datasets. Together, these form a comprehensive target system and quality control framework for the data foundation.

Guided by this strategic blueprint, Yazhouwan National Laboratory and Huawei jointly developed an AI data lake foundation built on Huawei's OceanStor Pacific all-flash scale-out storage. This platform aggregates agricultural research data previously scattered across China, publicly available datasets across the globe, enterprise-hosted data, and data collected by universities. Through a data classification and sharing mechanism—categorizing data into public, restricted, and confidential levels with corresponding security policies—the platform enables proper data circulation and controlled sharing, while achieving unified scheduling under a global file system.

More importantly, the data foundation is not only about integration, sharing, and utilization of existing data. To ensure data quality right at the source, both parties also established unified agricultural data collection standards, giving dispersed breeding data a "universal language". Whether generated in laboratories or collected in the field, all types of data can now be seamlessly integrated into the AI data lake.

With this data foundation in place, long-standing challenges such as fragmentation, inconsistency, and poor quality in agricultural datasets have been effectively addressed. National breeding data has become visible, manageable, and securely shareable. A high-quality breeding-oriented data corpus has been established, transforming previously isolated and difficult-to-share resources into reusable and circulatable national strategic assets.

The results are transformative. Driven by intelligent convergence and across-region data mobility, the era of siloed data storage and underutilization is officially over. A unified global data view, retrieval from exabyte-scale data in seconds, and on-demand sharing anytime, anywhere have become reality. Backed by the data foundation, Yazhouwan National Laboratory has successfully built a national precision breeding technology system, seamlessly empowering a massive network of 1 headquarters, 5 bases, N branches, and over 1,000 scientists.

From Experimental Fields to Data Bedrock: Revolutionizing the Future of Breeding

With the data foundation established, AI-driven breeding begins to accelerate.

For example, Yazhouwan National Laboratory has also developed an agricultural AI tool ecosystem based on Nexent, where intelligent agent systems seamlessly interwork with the data infrastructure, enabling automatic data recommendation and autonomous agent interaction.

The AI data lake provides strong, efficient support for a wide range of breeding AI tools. Researchers are no longer required to painstakingly locate specific datasets before they can begin their work, as was the case in the past. Looking ahead, Yazhouwan National Laboratory aims to deepen its collaboration with Huawei to build a multi-agent AI scientist system or an AI-driven breeding system that can serve both agricultural researchers and enterprise users.

In this sense, once the AI data foundation is established, solving the data challenge is only the beginning. The more profound transformation lies in the redefinition of traditional breeding workflows.

At its core, the AI data lake developed by Huawei is not a standalone product. It is deeply integrated with agricultural agent systems and breeding tool platforms, forming a closed-loop pipeline spanning data acquisition, analysis, decision-making, and execution.

In the past, when people thought of plant breeding, they often pictured researchers in the fields, working manually in muddy conditions. Today—and even more so in the future—breeding will increasingly become a fully modernized discipline. Researchers can use VR headsets, drones, and other front-end devices to capture field data in real time. This information is instantly fed into the AI data lake, where AI models generate breeding decisions that can be directly executed by robotic dogs and intelligent agricultural machinery. In this way, traditional weather-dependent and visual judgment–based breeding is upgraded into a precise, reliable, and intelligent paradigm.

From this perspective, the integration of Huawei AI data lake with bio-breeding is not merely a technological innovation. It represents a meaningful application of AI for All, demonstrating how technology can better serve national development and everyday livelihoods. It bridges the "last mile" between science and agriculture, transforming AI from a high-end laboratory capability into a practical tool rooted in the fields.

Moving forward, as this data bedrock continues to evolve, there is every reason to believe that China's seed industry will accelerate its efforts to close the technological gap, and develop more resilient, high-yield, and high-quality crop varieties adapted to diverse environments.

Arriba