← Latest papers
💬 NLP

Instruction-Guided Poetry Generation in Arabic and Its Dialects

This paper introduces a large-scale, instruction-based dataset in Modern Standard Arabic and its dialects that enables Large Language Models to controllably generate, revise, and analyze poetry according to specific user requirements like style and rhyme.

Original authors: Abdelrahman Sadallah, Kareem Elozeiri, Mervat Abassy, Rania Elbadry, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Fajri Koto

Published 2026-05-01
📖 4 min read☕ Coffee break read

Original authors: Abdelrahman Sadallah, Kareem Elozeiri, Mervat Abassy, Rania Elbadry, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Fajri Koto

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine Arabic poetry as a grand, ancient library filled with millions of beautiful books. For centuries, people have written poems in this library using strict rules about rhythm, rhyme, and style, much like building a house with specific blueprints. However, until now, computers (specifically Large Language Models or LLMs) have mostly been like librarians who can only read the books, count the pages, or guess who wrote them. They haven't been very good at writing new poems themselves, especially when asked to follow specific rules or write in different regional accents (dialects).

This paper introduces a new project called InstructPoet, which teaches computers how to become poets themselves. Here is how they did it, explained simply:

1. The Recipe Book (The Dataset)

To teach a computer to cook, you need a good cookbook. The researchers gathered a massive collection of Arabic poems from the internet, covering everything from ancient times to modern day.

  • The Cleanup: They organized this messy pile of poems into a neat, uniform format, like sorting a jumbled box of LEGOs into sorted bins.
  • The Labels: They added "metadata" (labels) to every poem, tagging them with details like "Romantic," "War," "Abbasid Era," or "Gulf Dialect."
  • The Translation: They didn't just use standard Arabic (MSA). They created instructions in five different flavors: Standard Arabic plus four major dialects (Gulf, Levantine, Nile Valley, and North African). This is like teaching the computer to speak not just formal English, but also Texan, Scottish, and Australian English, so it can talk to anyone naturally.

2. The Four Training Games (The Tasks)

Instead of just asking the computer to "write a poem," they set up four specific training games to make the computer smarter:

  • Generation (The Blank Page): "Write a romantic poem about the sea using a specific rhythm." The computer has to create something from scratch.
  • Continuation (The Cliffhanger): "Here are the first three lines of a poem; please finish the rest." The computer has to keep the story and rhythm going without breaking the flow.
  • Revision (The Editor): "Here is a poem with some typos and broken rhythm; fix it." The computer acts like a proofreader, restoring the poem to its perfect state.
  • Analysis (The Quiz): "Read this poem and tell me: What era is it from? What is the rhyme scheme?" The computer acts like a critic taking a test.

3. The Classroom (Training the Models)

The researchers took four different "student" computers (existing AI models) and taught them using this new dataset.

  • The Students: They used two models that are good at many languages (LLaMA and Qwen) and two that specialize in Arabic (ALLaM and Fanar).
  • The Teaching Style: They tried two methods. One was Joint Training (throwing all the games at the student at once), and the other was Curriculum Learning (teaching the easy stuff first, like analysis, and saving the hard stuff, like writing from scratch, for later).

4. The Report Card (Results)

After the training, they tested the students in two ways:

  • The Robot Grader: They used another AI to grade the poems on things like grammar, flow, and whether the computer actually followed the rules (like using the right rhyme).
  • The Human Judges: They hired native Arabic speakers who love poetry to read the poems and rate them on a scale of 1 to 5.

The Findings:

  • Huge Improvement: Before training, the computers were terrible at writing poetry that followed rules. After training, they got significantly better.
  • The Best Student: The ALLaM model, which was already designed for Arabic, became the top poet after training. It scored the highest in fluency and coherence.
  • The Hardest Task: Writing a poem from scratch (Generation) was the easiest for the computers. Fixing a broken poem (Revision) was the hardest, because it required the computer to understand the "soul" of the poem while fixing the technical errors.
  • Dialects Worked: The computers got much better at understanding and writing in different Arabic dialects, not just the formal version.

The Bottom Line

This paper doesn't claim that computers will replace human poets or that they can write perfect, soul-stirring masterpieces yet. Instead, it shows that we can now build a digital assistant that helps people write poetry. If you want a poem in a specific style, with a specific rhyme, in your local dialect, this new system can help you draft it, fix your mistakes, or even analyze what you've written. It turns the computer from a passive reader into an active, rule-following creative partner.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →