← Latest papers
💬 NLP

ERNIE 5.0 Technical Report

This paper introduces ERNIE 5.0, the first publicly disclosed production-scale trillion-parameter unified autoregressive foundation model that natively supports multimodal understanding and generation across text, image, video, and audio through an ultra-sparse mixture-of-experts architecture and a novel elastic training paradigm.

Original authors: Haifeng Wang, Hua Wu, Tian Wu, Yu Sun, Jing Liu, Dianhai Yu, Yanjun Ma, Jingzhou He, Zhongjun He, Dou Hong, Qiwen Liu, Shuohuan Wang, Junyuan Shang, Zhenyu Zhang, Yuchen Ding, Jinle Zeng, Jiabin Yang
Published 2026-02-05
📖 5 min read🧠 Deep dive

Original authors: Haifeng Wang, Hua Wu, Tian Wu, Yu Sun, Jing Liu, Dianhai Yu, Yanjun Ma, Jingzhou He, Zhongjun He, Dou Hong, Qiwen Liu, Shuohuan Wang, Junyuan Shang, Zhenyu Zhang, Yuchen Ding, Jinle Zeng, Jiabin Yang, Liang Shen, Ruibiao Chen, Weichong Yin, Siyu Ding, Dai Dai, Shikun Feng, Siqi Bao, Bolei He, Yan Chen, Zhenyu Jiao, Ruiqing Zhang, Zeyu Chen, Qingqing Dang, Kaipeng Deng, Jiajun Jiang, Enlei Gong, Guoxia Wang, Yanlin Sha, Yi Liu, Yehan Zheng, Weijian Xu, Jiaxiang Liu, Zengfeng Zeng, Yingqi Qu, Zhongli Li, Zhengkun Zhang, Xiyang Wang, Zixiang Xu, Xinchao Xu, Zhengjie Huang, Dong Wang, Bingjin Chen, Yue Chang, Xing Yuan, Shiwei Huang, Qiao Zhao, Xinzhe Ding, Shuangshuang Qiao, Baoshan Yang, Bihong Tang, Bin Li, Bingquan Wang, Binhan Tang, Binxiong Zheng, Bo Cui, Bo Ke, Bo Zhang, Bowen Zhang, Boyan Zhang, Boyang Liu, Caiji Zhang, Can Li, Chang Xu, Chao Pang, Chao Zhang, Chaoyi Yuan, Chen Chen, Cheng Cui, Chenlin Yin, Chun Gan, Chunguang Chai, Chuyu Fang, Cuiyun Han, Dan Zhang, Danlei Feng, Danxiang Zhu, Dong Sun, Dongbo Li, Dongdong Li, Dongdong Liu, Dongxue Liu, Fan Ding, Fan Hu, Fan Li, Fan Mo, Feisheng Wu, Fengwei Liu, Gangqiang Hu, Gaofeng Lu, Gaopeng Yong, Gexiao Tian, Guan Wang, Guangchen Ni, Guangshuo Wu, Guanzhong Wang, Guihua Liu, Guishun Li, Haibin Li, Haijian Liang, Haipeng Ming, Haisu Wang, Haiyang Lu, Haiye Lin, Han Zhou, Hangting Lou, Hanwen Du, Hanzhi Zhang, Hao Chen, Hao Du, Hao Liu, Hao Zhou, Haochen Jiang, Haodong Tian, Haoshuang Wang, Haozhe Geng, Heju Yin, Hong Chen, Hongchen Xue, Hongen Liu, Honggeng Zhang, Hongji Xu, Hongwei Chen, Hongyang Zhang, Hongyuan Zhang, Hua Lu, Huan Chen, Huan Wang, Huang He, Hui Liu, Hui Zhong, Huibin Ruan, Jiafeng Lu, Jiage Liang, Jiahao Hu, Jiahao Hu, Jiajie Yang, Jialin Li, Jian Chen, Jian Wu, Jianfeng Yang, Jianguang Jiang, Jianhua Wang, Jianye Chen, Jiaodi Liu, Jiarui Zhou, Jiawei Lv, Jiaxin Zhou, Jiaxuan Liu, Jie Han, Jie Sun, Jiefan Fang, Jihan Liu, Jihua Liu, Jing Hu, Jing Qian, Jing Yan, Jingdong Du, Jingdong Wang, Jingjing Wu, Jingyong Li, Jinheng Wang, Jinjin Li, Jinliang Lu, Jinlin Yu, Jinnan Liu, Jixiang Feng, Jiyi Huang, Jiyuan Zhang, Jun Liang, Jun Xia, Jun Yu, Junda Chen, Junhao Feng, Junhong Xiang, Junliang Li, Kai Liu, Kailun Chen, Kairan Su, Kang Hu, Kangkang Zhou, Ke Chen, Ke Wei, Kui Huang, Kun Wu, Kunbin Chen, Lei Han, Lei Sun, Lei Wen, Linghui Meng, Linhao Yu, Liping Ouyang, Liwen Zhang, Longbin Ji, Longzhi Wang, Meng Sun, Meng Tian, Mengfei Li, Mengqi Zeng, Mengyu Zhang, Ming Hong, Mingcheng Zhou, Mingming Huang, Mingxin Chen, Mingzhu Cai, Naibin Gu, Nemin Qiu, Nian Wang, Peng Qiu, Peng Zhao, Pengyu Zou, Qi Wang, Qi Xin, Qian Wang, Qiang Zhu, Qianhui Luo, Qianwei Yang, Qianyue He, Qifei Wu, Qinrui Li, Qiwen Bao, Quan Zhang, Quanxiang Liu, Qunyi Xie, Rongrui Zhan, Rufeng Dai, Rui Peng, Ruian Liu, Ruihao Xu, Ruijie Wang, Ruixi Zhang, Ruixuan Liu, Runsheng Shi, Ruting Wang, Senbo Kang, Shan Lu, Shaofei Yu, Shaotian Gong, Shenwei Hu, Shifeng Zheng, Shihao Guo, Shilong Fan, Shiqin Liu, Shiwei Gu, Shixi Zhang, Shuai Yao, Shuang Zhang, Shuangqiao Liu, Shuhao Liang, Shuwei He, Shuwen Yang, Sijun He, Siming Dai, Siming Wu, Siyi Long, Songhe Deng, Suhui Dong, Suyin Liang, Teng Hu, Tianchan Xu, Tianliang Lv, Tianmeng Yang, Tianyi Wei, Tiezhu Gao, Ting Sun, Ting Zhang, Tingdan Luo, Wei He, Wei Luan, Wei Yin, Wei Zhang, Wei Zhou, Weibao Gong, Weibin Li, Weicheng Huang, Weichong Dang, Weiguo Zhu, Weilong Zhang, Weiqi Tan, Wen Huang, Wenbin Chang, Wenjing Du, Wenlong Miao, Wenpei Luo, Wenquan Wu, Xi Shi, Xi Zhao, Xiang Gao, Xiangguo Zhang, Xiangrui Yu, Xiangsen Wang, Xiangzhe Wang, Xianlong Luo, Xianying Ma, Xiao Tan, Xiaocong Lin, Xiaofei Wang, Xiaofeng Peng, Xiaofeng Wu, Xiaojian Xu, Xiaolan Yuan, Xiaopeng Cui, Xiaotian Han, Xiaoxiong Liu, Xiaoxu Fei, Xiaoxuan Wu, Xiaoyu Wang, Xiaoyu Zhang, Xin Sun, Xin Wang, Xinhui Huang, Xinming Zhu, Xintong Yu, Xinyi Xu, Xinyu Wang, Xiuxian Li, XuanShi Zhu, Xue Xu, Xueying Lv, Xuhong Li, Xulong Wei, Xuyi Chen, Yabing Shi, Yafeng Wang, Yamei Li, Yan Liu, Yanfu Cheng, Yang Gao, Yang Liang, Yang Wang, Yang Wang, Yang Yang, Yanlong Liu, Yannian Fu, Yanpeng Wang, Yanzheng Lin, Yao Chen, Yaozong Shen, Yaqian Han, Yehua Yang, Yekun Chai, Yesong Wang, Yi Song, Yichen Zhang, Yifei Wang, Yifeng Guo, Yifeng Kou, Yilong Chen, Yilong Guo, Yiming Wang, Ying Chen, Ying Wang, Yingsheng Wu, Yingzhan Lin, Yinqi Yang, Yiran Xing, Yishu Lei, Yixiang Tu, Yiyan Chen, Yong Zhang, Yonghua Li, Yongqiang Ma, Yongxing Dai, Yongyue Zhang, Yu Ran, Yu Sun, Yu-Wen Michael Zhang, Yuang Liu, Yuanle Liu, Yuanyuan Zhou, Yubo Zhang, Yuchen Han, Yucheng Wang, Yude Gao, Yuedong Luo, Yuehu Dong, Yufeng Hu, Yuhui Cao, Yuhui Yun, Yukun Chen, Yukun Gao, Yukun Li, Yumeng Zhang, Yun Fan, Yun Ma, Yunfei Zhang, Yunshen Xie, Yuping Xu, Yuqin Zhang, Yuqing Liu, Yurui Li, Yuwen Wang, Yuxiang Lu, Zefeng Cai, Zelin Zhao, Zelun Zhang, Zenan Lin, Zezhao Dong, Zhaowu Pan, Zhaoyu Liu, Zhe Dong, Zhe Zhang, Zhen Zhang, Zhengfan Wu, Zhengrui Wei, Zhengsheng Ning, Zhenxing Li, Zhenyu Li, Zhenyu Qian, Zhenyun Li, Zhi Li, Zhichao Chen, Zhicheng Dong, Zhida Feng, Zhifan Feng, Zhihao Deng, Zhijin Yu, Zhiyang Chen, Zhonghui Zheng, Zhuangzhuang Guo, Zhujun Zhang, Zhuo Sun, Zichang Liu, Zihan Lin, Zihao Huang, Zihe Zhu, Ziheng Zhao, Ziping Chen, Zixuan Zhu, Ziyang Xu, Ziyi Liang, Ziyuan Gao

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: The "Swiss Army Knife" Brain

Imagine most AI models today are like a team of specialists: one person writes, another draws, and a third sings. They might talk to each other, but they are separate people with separate brains.

ERNIE 5.0 is different. It is a single, massive brain that learns to write, draw, sing, and make videos all at the same time, from day one. Instead of adding a "drawing module" to a "writing brain," ERNIE 5.0 is trained from scratch to understand that a picture, a sound, and a word are all just different types of "tokens" (like puzzle pieces) that fit into the same big puzzle.

1. The Engine: A Smart, Sparse Factory

The paper describes the model's architecture as an Ultra-Sparse Mixture-of-Experts (MoE).

  • The Analogy: Imagine a massive factory with 1,000 different specialized workers (experts). When a task comes in, the factory manager (the router) doesn't wake up all 1,000 workers. Instead, they only wake up the top 2 or 3 workers who are best at that specific job.
  • The Innovation: In older models, the manager might have different rules for "writing tasks" vs. "drawing tasks." In ERNIE 5.0, the manager is modality-agnostic. It doesn't care if the request is a poem or a painting; it just looks at the content and picks the best workers. This allows the model to be huge (trillions of parameters) but still fast and efficient because it only uses a tiny fraction of its brain for any single task.

2. The Learning Style: "One Size Fits All" (Elastic Training)

Usually, if a company wants a small AI for a phone and a big AI for a server, they have to train two different models or shrink the big one later (like cutting a cake). This is wasteful and often ruins the cake's taste.

ERNIE 5.0 uses Elastic Training.

  • The Analogy: Imagine training a gymnast. Instead of just training them to do a full routine, you train them to do the full routine, but also randomly ask them to skip a few flips or use a shorter beam during practice.
  • The Result: By the end of training, this gymnast is so well-prepared that they can perform the full routine perfectly, or they can instantly switch to a shorter, simpler routine without needing extra practice. This means one single ERNIE 5.0 model can be "shrunk" on the fly to fit a phone, a tablet, or a supercomputer without losing much performance.

3. Seeing and Hearing: Speaking the Same Language

To handle images and audio, the model uses special translators (tokenizers) to turn pictures and sounds into the same "language" as text.

  • Vision (Images/Video): The model treats a single image as a "one-frame video." It learns to predict the next "scale" of detail (like zooming in) and the next frame in time. It uses a "hybrid" approach, looking at the big picture (semantics) and the tiny details (pixels) simultaneously, like an artist who understands both the story of a painting and the brushstrokes.
  • Audio (Speech/Sound): It breaks sound down into layers, like peeling an onion. The first layer captures the "meaning" (what is being said), and the deeper layers capture the "texture" (the voice tone, the background noise). It predicts these layers one by one, from the big idea down to the fine details.

4. Getting Smarter: Reinforcement Learning with Hints

After the initial training, the model needs to learn how to reason and follow complex instructions. This is done through Reinforcement Learning (RL).

  • The Challenge: Sometimes the model gets stuck on hard problems and stops trying (entropy collapse).
  • The Solution: The researchers introduced Adaptive Hint-based Learning.
    • The Analogy: Imagine a student taking a hard math test. If they get stuck, the teacher doesn't just give them the answer. Instead, the teacher gives a tiny hint ("Think about the first step"). As the student gets better, the teacher gives fewer hints until the student can solve it alone.
    • This helps the model learn difficult tasks without getting frustrated or giving up.

5. The Results: Balanced and Ready for the Real World

The paper claims that ERNIE 5.0 is the first production-scale model of its kind (trillion parameters) that natively handles text, images, video, and audio together.

  • Performance: It performs as well as (or better than) specialized models in reading, math, coding, and drawing.
  • Efficiency: Because of the "elastic" design, you can turn it down to use only 35% of its total power and still get 95% of the results. This makes it practical to run on different devices, from powerful servers to smaller gadgets.

In summary: ERNIE 5.0 is a unified, flexible, and efficient AI brain that learns everything at once. It doesn't need to be retrained to fit different devices, and it treats pictures, sounds, and words as part of the same conversation, rather than separate subjects.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →