← Latest papers
🤖 AI

Phoenix-VL 1.5 Medium Technical Report

Phoenix-VL 1.5 Medium is a 123B-parameter natively multimodal and multilingual foundation model specifically adapted for the Singapore context through extensive localized pretraining and alignment, achieving state-of-the-art performance on regional benchmarks while maintaining global competitiveness in general intelligence.

Original authors: Team Phoenix, :, Arka Ray, Askar Ali Mohamed Jawad, Biondi Lee, Elijah Seah, Eva Lim, Fiona Teo, Grace Toh, Guang Xiang Teo, Jun En Tan, Jia Hui Bong, Jiale Wang, Jonathan Ng, Justin Tan, Kai Zhe Yew
Published 2026-05-12
📖 5 min read🧠 Deep dive

Original authors: Team Phoenix, :, Arka Ray, Askar Ali Mohamed Jawad, Biondi Lee, Elijah Seah, Eva Lim, Fiona Teo, Grace Toh, Guang Xiang Teo, Jun En Tan, Jia Hui Bong, Jiale Wang, Jonathan Ng, Justin Tan, Kai Zhe Yew, Matthew Ong, Shun Yi Yeo, Wen Jett Lam, Wen Xiu Tan, Ze Yu Zhang, Gee Wah Ng, Chee Wee Ang, Mistral AI, :, Adrien Sadé, Guillaume Kunsch, Jia Sin Loh, Nicolas Schuhl, Rupert Menneer, Umar Jamil, Vincent Maladière, Yimu Pan

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

The Big Idea: A "Local Expert" Super-Brain

Imagine you have a brilliant, world-traveled professor who knows everything about science, math, and global history. This is the Mistral Medium 3.1 model the researchers started with. It's incredibly smart, but if you ask it about a specific local law in Singapore or the name of a famous playground in the Toa Payoh neighborhood, it might guess wrong or give a generic answer.

Phoenix-VL 1.5 Medium is what happens when you take that world-traveled professor and send them to a rigorous, 12-month "Singapore Immersion Bootcamp." The goal was to create a Sovereign AI—a digital brain that belongs to Singapore, understands its unique culture and laws, and can work securely without needing to look things up on the internet (like an "air-gapped" system that doesn't talk to the outside world).

The Training Recipe: How They Did It

The researchers didn't just dump a pile of books on the model. They used a specific, four-step cooking recipe to bake this new AI:

  1. The Base Layer (Continued Pretraining):
    Think of this as feeding the model a massive, specialized library. They gave it 1 trillion tokens (a token is like a word or a piece of a word) of data. This wasn't just random internet text; it was a curated mix heavy on Singaporean laws, regional languages (like Malay, Tamil, and Chinese), and local culture.

    • The Analogy: Imagine the professor reading every Singaporean newspaper, government gazette, and local novel for a year, while still keeping their global knowledge intact.
  2. The Long-Read Extension (Long Context):
    They taught the model to remember huge amounts of information at once. They extended its "short-term memory" to hold 131,072 tokens (roughly the size of a thick novel).

    • The Analogy: The professor can now read an entire 500-page legal document in one sitting and remember the details from page 1 when they get to page 400, without forgetting the beginning.
  3. The Fine-Tuning (Instruction & Domain Training):
    This is where the model learned how to behave. They showed it specific examples of how to answer questions about Singaporean government policies and how to describe local scenes (like a "Toa Payoh Dragon Playground" instead of just a generic "playground").

    • The Analogy: The professor is now taking a "Customer Service and Local Etiquette" class. They learned that when a Singaporean asks about a law, they need a precise answer, not a vague guess. They also learned to look at pictures of local landmarks and describe them accurately.
  4. The Polish (Online Direct Preference Optimization):
    Finally, human teachers reviewed the model's answers and ranked them. If the model was too chatty, hallucinated (made things up), or was rude, the teachers corrected it.

    • The Analogy: A final "dress rehearsal" where the professor practices answering questions, and a panel of judges gives them feedback to ensure they are helpful, honest, and safe before they go on stage.

What Makes It Special?

The paper highlights three main superpowers of this new model:

  • It's a Local Legend: On tests specifically designed for Singapore (like knowing the names of all 16 government ministries or understanding local safety laws), Phoenix-VL 1.5 Medium beat much larger, more famous models.
    • The Result: It scored higher than models with 400 billion parameters (like Llama 4 Maverick) on local knowledge. This proves that deep local training is more important than just having a bigger brain.
  • It Didn't Forget the World: A common fear is that teaching a model about one specific topic makes it forget everything else. The paper claims Phoenix-VL 1.5 Medium kept its general smarts. It is still great at math, coding, and general reasoning, just as good as other top-tier models.
  • It Sees and Understands: Unlike older models that only read text, this one is "multimodal." It can look at a picture of a Singaporean street scene or a complex chart and explain it in the local context.

Safety: The "Honesty" Filter

The researchers built a special safety framework tailored to Singapore's rules.

  • The "Don't Make Things Up" Rule: In legal or government contexts, making up a fact is dangerous. The model was trained to say "I don't know" rather than guessing.
  • The "No Bad Actors" Rule: The model is resistant to "jailbreaks" (tricks people use to make AI do bad things) and won't leak its own secret instructions.
  • The Analogy: Think of it as a very disciplined civil servant who knows exactly what they are allowed to say and what they must refuse to answer, ensuring they never accidentally break the law or spread misinformation.

The Bottom Line

Phoenix-VL 1.5 Medium is a 123-billion-parameter AI model that proves you don't need the biggest model in the world to be the smartest local expert. By carefully curating data and training specifically for Singapore's unique laws, languages, and culture, the creators built a model that is:

  1. Deeply knowledgeable about Singapore (better than giants).
  2. Globally competitive in general smarts.
  3. Safe and secure for use in government and private sectors without needing the internet.

It's a "sovereign asset," meaning it's a digital tool built specifically for Singapore's needs, owned and controlled by the nation.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →