← Latest papers
💬 NLP

Gemma 4 Technical Report

The Gemma 4 Technical Report introduces a new generation of open-weight, natively multimodal models featuring diverse architectures, an encoder-free design for raw audio and image processing, and a thinking mode, all of which collectively deliver significant advancements in efficiency, reasoning, and performance across STEM and long-context benchmarks.

Original authors: Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Victor Cotruta, Alice Couck
Published 2026-07-07
📖 4 min read☕ Coffee break read

Original authors: Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Clément Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst, Jiaxian Guo, Cassidy Hardin, Yanzhang He, Steven M. Hernandez, Omri Homburger, Léonard Hussenot, Juyeong Ji, Armand Joulin, Aishwarya Kamath, Parnian Kassraie, Olivier Lacombe, Preethi Lahoti, Gaël Liu, Gus Martins, Luciano Martins, Tatiana Matejovicova, Ramona Merhej, Nikola Momchev, Sneha Mondal, Ryan Mullins, Sindhu Raghuram Panyam, Shreya Pathak, Sarah Perrin, André Susano Pinto, Etienne Pot, Angéline Pouget, Alexandre Ramé, Sabela Ramos, Douglas Reid, David Rim, Morgane Rivière, Karsten Roth, Louis Rouillard, Omar Sanseviero, Pier Giuseppe Sessa, Shane Settle, Danila Sinopalnikov, Sara Smoot, Piotr Stanczyk, Andreas Steiner, Lawrence Stewart, Ilya Tolstikhin, Michael Tschannen, Anton Tsitsulin, Nino Vieillard, Renjie Wu, Pingmei Xu, Haichuan Yang, Edouard Yvinec, Li Zhang, Joe Zou, Nicolas Aagnes, Abdelrahman Abdelhamed, Shivani Agrawal, Shubham Agrawal, Ibrahim Alabdulmohsin, Jean Baptiste Alayrac, Uri Alon, Chandramouli Amarnath, Ankesh Anand, Chrysovalantis Anastasiou, Setareh Ariafar, François-Xavier Aubet, Kyriakos Axiotis, Federico Barbero, Joelle Barral, Alexei Bendebury, Urs Bergmann, Stanley Bileschi, Kat Black, Mathieu Blondel, Sebastian Borgeaud, Arthur Bražinskas, Ryan Burnell, Robert Busa-Fekete, Mu Cai, Glenn Cameron, Charlotte Caucheteux, Garima Chadha, Jetha Chan, Aditya Chawla, Blake Jianhang Chen, Jesse Chen, Lin Chen, Xu Chen, Derek Cheng, Tzu-hsiang Chien, Nikolai Chinaev, Yi Chou, Zhaohui Chu, Benjamin Coleman, Pooja Consul, Sam Conway-Rahman, Scott Crowell, Dylan Cutler, Vivek Dani, Samira Daruki, Anil Das, Daniel Deutsch, Nishanth Dikkala, Li Ding, Qiuhan Ding, Shenil Dodhia, Konstantin Donhauser, Tulsee Doshi, Anca Dragan, Alex Druinsky, Sahil Dua, Zoltan Egyed, Danielle Eisenbud, Daniel Eppens, Cindy Fan, Bahare Fatemi, Yassir Fathullah, Vlad Feinberg, Milen Ferev, Takumi Fujimoto, Isaac Galatzer-Levy, João Gante, Simon Geisler, Soham Ghosal, Antonious M. Girgis, Alec Go, Alhaad Gokhale, Alex Grills, Yiming Gu, Pramod Gupta, Guru Guruganesh, Raia Hadsell, Hamza Harkous, Jitendra Harlalka, Demis Hassabis, Anja Hauth, Joe Heyward, Arian Hosseini, Chih-Yang Hsia, I-Hung Hsu, Xiaopeng Huang, Yangsibo Huang, Kevin Hui, Adrian Hutter, Te I, Fotis Iliopoulos, Advait Jain, Ganesh Jawahar, Ziwei Ji, Qilin Jin, Melvin Johnson, Kandarp Joshi, Arun Kandoor, Wang-Cheng Kang, Koray Kavukcuoglu, Mehran Kazemi, Kathleen Kenealy, Amr Khalifa, Phoebe Kirk, Suraj Kothawade, Vitaly Kovalev, Neel Kovelamudi, Adam Kraft, Ravin Kumar, Harish Kuppam, Justin Lannin, Chen-Yu Lee, Seungji Lee, Dmitry Lepikhin, Dongdong Li, Qiujia Li, Valentin Liévin, Ethan Lin, Ziqian Lin, Casper Liu, Tianlin Liu, Tianqi Liu, Xin Liu, Mayank Lunayach, Min Ma, Gagan Madan, Andrii Maksai, Eric Malmi, Michal Matuszak, Daniel McDuff, Gaurav Menghani, Daniil Mirylenka, Karolis Misiunas, Vedant Misra, Andreea Mitran, Kareem Mohamed, Maksim Mukha, Eric Noland, James O'Donnell, Kate Olszewska, Bernett Orlando, Wanqiong Pan, Rina Panigrahy, Unnati Parekh, Chunjong Park, Eric Paskie, Liqian Peng, Bryce Petrini, Slav Petrov, Jonas Pfeiffer, Bilal Piot, Martyna Plomecka, Siim Poder, Octavio Ponce, Arijit Pramanik, David Racz, Anish Rajan, Michelle Ramanovich, Anand Rao, Marvin Ritter, Vitor Rodrigues, Evan Rosen, Mikołaj Rybiński, Noveen Sachdeva, Michaël E. Sander, Rohit Sathyanarayana, Sagar Savla, Samuel Schmidgall, Tal Schuster, Benoit Seguin, Andrew Sellergren, Aliaksei Severyn, Izhak Shafran, Dhruv Shah, Yuan Shangguan, Ashish Shenoy, Pradeep Shenoy, Rakesh Shivanna, Pauline Sho, Lucas Spangher, Wojciech Stokowiec, Tim Strother, Yao Su, Yinghao Sun, Mukund Sundararajan, Andrea Tacchetti, Mor Hazan Taege, Pouya Tafti, Chetan Tekur, Rahul Thapa, Madeleine Traverse, Lenart Treven, Tao Tu, Chien Te Tung, Petar Veličković, Malini Pooni Venkat, Sagar Gubbi Venkatesh, Vidya Venkiteswaran, Francesco Visin, Alex Vitvitskyi, Kiran Vodrahalli, Weiyi Wang, Xin Wang, Tris Warkentin, Jan Wassenberg, John Wieting, Lechao Xiao, Hao Xu, Yuhui Xu, Fuzhao Xue, Arun Yadav, Jun Yan, Antoine Yang, Lin Yang, Ming-Hsuan Yang, Ziyu Ying, Jae Hyeon Yoo, Sajjad Zafar, Fred Zhang, Jiageng Zhang, Jianyi Zhang, Xiaofan Zhang, Chao Zhao, David Zhou, Chen Zou

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine Google DeepMind has just unveiled Gemma 4, a new family of "smart assistants" that are open for anyone to use, tweak, and build upon. Think of these not as single, giant robots, but as a whole toolbox of different-sized brains, ranging from a tiny, pocket-sized helper to a massive, super-computer-level thinker.

Here is a breakdown of what makes Gemma 4 special, using simple analogies:

1. A Toolbox of Different Sizes

Previously, you might have had one big brain or one small brain. Gemma 4 offers a whole set:

  • The Pocket Helpers (2.3B and 4.5B): These are like smart watches. They are small enough to run on your phone or laptop without needing a massive server farm.
  • The Workhorses (12B and 31B): These are like powerful desktop computers, capable of handling complex tasks.
  • The Specialist (26B-A4B): This is a "Mixture of Experts" model. Imagine a team of 26 people where, for any given question, only the 4 most relevant experts step forward to answer. This makes it incredibly fast and efficient while still being very smart.

2. The "Thinking Cap" (Reasoning Mode)

One of the biggest upgrades is a new "Thinking Mode."

  • Before: If you asked a model a hard math problem, it would guess the answer immediately.
  • Now: The model puts on a "thinking cap." It whispers its thought process to itself first (like a student working out a problem on scratch paper) before writing down the final answer. This allows it to solve tricky math and coding problems much better, just like a human does when they take their time to think.

3. Seeing and Hearing Without Glasses or Ears

Older models needed special "glasses" (vision encoders) and "ears" (audio encoders) to understand pictures and sounds. These were heavy, separate devices attached to the brain.

  • The New Approach: The 12B model is encoder-free. It's like the brain itself has learned to see and hear directly. Instead of wearing heavy glasses, it looks at raw pixels of an image or raw sound waves of a voice and understands them instantly. This makes the whole system lighter, faster, and less cluttered.

4. The Infinite Library (Long Context)

Imagine trying to read a 1,000-page book and remembering every detail. Old models would get "memory fog" and forget the beginning by the time they reached the end.

  • The Fix: Gemma 4 uses a clever memory trick. It treats the book like a sliding window: it remembers the last few pages in high definition (local attention) but keeps a summarized, efficient index of the whole book (global attention).
  • The Result: It can read and remember massive amounts of text (up to 128k or 256k tokens) without running out of memory, and it does this using 37.5% less memory than before.

5. Speeding Up the Race (Efficiency)

  • Predicting the Future: The models use a "drafter" head. Imagine a race car driver who not only drives the car but also predicts the next three turns before they happen. This allows the model to guess the next few words in a sentence instantly, making it speak much faster.
  • Compressing the Brain: The team trained the models to be "quantized." Think of this like packing a suitcase. Instead of packing heavy, bulky clothes (full precision), they fold everything tightly (compressed weights) so it fits in a smaller bag. You can run these models on mobile devices with almost no loss in quality.

6. How Smart Are They?

The paper compares Gemma 4 to other top-tier models (like Claude or DeepSeek) in a "blind taste test" called the Arena.

  • The Result: The 31B model is the top-ranked open model (meaning anyone can download it) in its size category. It performs just as well as models that are 10 times larger.
  • The Small Giant: The tiny 2.3B model performs as well as the previous generation's 27B model, meaning it's 10 times more efficient than before.

7. Safety and Responsibility

Just like you wouldn't let a child drive a car without training, the team built safety into Gemma 4 from the ground up.

  • They filtered out dangerous or harmful data before training.
  • They tested the models rigorously to ensure they don't generate hate speech, dangerous instructions, or harmful content.
  • They acknowledge that while these tools are powerful, they must be used responsibly, and they provide guidelines to help developers keep things safe.

In short: Gemma 4 is a new generation of open AI that is faster, smarter at reasoning, capable of seeing and hearing directly, and able to remember huge amounts of information—all while being small enough to run on your own devices.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →