author:gcc
HiFloat Community Officially Launched, Accelerating the Formation of the Low-Precision Computing Industrial Ecosystem
On August 26, the “Low-Precision Computing Technology Seminar and HiFloat Community Launch Event” was held in Beijing. The event covered low-bit-width data format standards, hardware-software co-design, HiFloat technologies and applications, Ascend parallel acceleration, efficient data transmission, and model quantization. The GCC-HiFloat
Community was officially launched, and a series of the community’s interim achievements were released. Experts from research institutions, universities, chip companies, and artificial intelligence companies gathered to discuss the development path for low-precision computing as it moves from technological innovation toward standardization and industrialization.
Huang Huanqing, Deputy Secretary-General of the GCC Intelligent Computing Industry Development Committee, introduced the overall plan for the HiFloat Community. Built on the principles of openness, sharing, and joint development, the community will open up HiFloat-related technologies and standards, share research and practical results on low-bit-width data formats, and attract partners across the industrial chain to participate. The community will continue to deliver outputs around standards, hardware and software framework support, mainstream model validation, the developer ecosystem, toolchains, and industry research, while driving upstream ecosystem support and participating in thedevelopment of national standards and major standardization venues such as IEEE.
The event announced the membership of the HiFloat Community Technical Committee, which includes: China Electronics Standardization Institute (CESI), Shanghai AI Laboratory, Huawei Technologies Co., Ltd., South China University of Technology, Nanjing University, Qualcomm Wireless Communication Technologies (China) Co., Ltd., iFLYTEK Co., Ltd., Shanghai SenseTime Technology Development Co., Ltd., Approaching.ai, Beijing Digital China Kuntai Information Technology Co., Ltd., and MetaX Integrated Circuits (Shanghai) Co., Ltd.
The event also released three interim advances of the HiFloat Community: the association standard HiFloat8 Data Format Technical Specification, the call for proposals for the HiFloat4 Data Format Technical Specification association standard, and the results of the IEEE ICME Low-Bit-Width Large Model Quantization Challenge.


Zhang Qi, an AI standardization expert at China Electronics Standardization Institute (CESI), delivered a keynote titled “Progress and Planning of National Standards for Low-Bit-Width Data Formats.” He introduced that the national standard Technical Specification for AI Low-Bit-Width Floating-Point Data Formats is currently under development. In the next phase, training and inference validation on large models will be used to guide industrial partners such as chip and model companies in carrying out pilots, and work on international standards will be further advanced.
Lu Jinming, Assistant Professor and Distinguished Research Fellow at Nanjing University, explained why low-precision computing is worth pursuing from the perspective of hardware-software co-design. He introduced the evolution of low-bit-width data formats and block floating-point (BFP) formats, and discussed hierarchical design using formats such as HiFloat4 as examples. At extremely low bit widths, selecting the appropriate precision, scaling granularity, rounding mode, and accumulation precision for different data requires the co-design of algorithms, compilers, operators, chip microarchitecture, and even storage and communication systems.
Luo Yuanyong, a technical expert at Huawei Technologies Co., Ltd., further unpacked the design principles and philosophy of the HiFloat data format, as well as its usage in training and inference. Low precision is not simply about turning a model into low bit with one click. Different weights, activations, gradients, and sensitivity layers do not tolerate precision in the same way; targeted optimization is required through methods such as scale configuration, precision allocation, rounding strategies, and operator fusion.
Li Sitan, a processor algorithm expert, pushed the question further: “What happens after it is actually put into models?” She demonstrated on site the practical applications and performance optimizations of the HiFloat data format on models including LongCat, DeepSeek, Wan, and Qwen, covering model training, inference, and quantization optimization. This further shows that HiFloat is evolving from a data format proposal into a technical system that can be validated in practice across different models and tasks.
Lu Lu, Professor at South China University of Technology, shared insights on the performance optimization of related operators and template libraries, as well as industrial ecosystem practice. Taking some general-purpose computing fusion tests as examples, by partitioning data and optimizing pipelined concurrency, double-digit percentage performance improvements over baselines were achieved in certain scenarios. Related results continue to be extended into practical applications through open-source template libraries, compilation tools, and industry projects. For low-precision computing to truly deliver value, it must ultimately come down to the co-optimization of chips, operators, compilers, and specific applications.
Ren Feng, a technical expert at Approaching.ai and author of the Mooncake Transfer Engine, shifted the perspective from “computing” to “data transmission.” He pointed out that low-precision inference reduces memory footprint and makes computing and memory access faster; however, data loading, KV Cache movement, and cross-node communication then account for a larger share of overall latency.Low precision makes each piece of data smaller, but it does not automatically solve the questions of where data should be placed, which link it should travel through, or how to make full use of multiple network interface cards. Computing optimization and transmission optimization must therefore advance in parallel.
Zhou Rongchen, a senior Ascend engineer at Huawei Technologies Co., Ltd., introduced the quantization tool msModelSlim, which engineers the manual debugging process in quantization and describes different quantization algorithms through composable configurations, enabling one solution to be conveniently reused across different models.Mechanisms such as quantization modes and model adaptation protocols also connect tools, operators, and inference frameworks. In the future, the tool will incorporate a quantization knowledge base and Agent capabilities, driving quantization schemes, sensitivity layer handling, and model adaptation further toward automation.
The event brought together industrial chain partners from models, algorithms, chips, software, networking, and the standards system. The atmosphere at the venue was lively, with participating guests engaging in thorough discussion and in-depth exchanges.
Low-precision computing is not an engineering effort that can be completed by any single company, single model, or single data format. Only when data formats can be efficiently executed by chips, natively supported by software frameworks, fully validated across different models, and underpinned by open standards and a developer ecosystem, can the performance and energy-efficiency advantages brought by low precision be truly translated into industrial value.
Going forward, the HiFloat Community will continue to uphold the principles of openness, sharing, and joint development, connect with more industrial partners, and jointly explore the higher efficiency of AI computing at “lower precision.”
If you would like to learn more about or participate in the GCC-HiFloat Community’s upcoming technical seminars, standards co-development, industry co-innovation, and ecosystem cooperation activities, please contact: gcc_hifloat@gccorg.com
The English website is now being developed.
We sincerely appreciate your attention
and continued support.