Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
This study demonstrates that synthetic rewriting acts as a quality multiplier rather than a substitute for data curation in Portuguese continued pretraining, significantly boosting performance only when applied to high-quality source data and primarily at larger model scales.
Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Large language models are the engines behind many of the most advanced artificial intelligence systems today, capable of writing text, answering questions, and solving problems with a fluency that often mimics human thought. These systems learn by reading vast amounts of text from the internet, absorbing patterns in how words are used and how ideas are connected. However, this learning process is not equally effective for every language. While English has a massive library of high-quality text available for training, many other languages, such as Portuguese, have far fewer resources. To bridge this gap, researchers often take a model trained primarily on English and continue teaching it using smaller amounts of text in the target language. A growing question in the field is whether we can make this limited data even more powerful by using artificial intelligence to rewrite it. The idea is that if a computer can take a messy or simple sentence and rewrite it to be clearer, more structured, or more detailed, the model might learn better from the improved version. But a crucial uncertainty remains: does this rewriting technique work equally well on poor-quality text, or does it only amplify the value of text that was already good to begin with?
A team of researchers at the University of Campinas in Brazil set out to answer this question by conducting a controlled experiment with Portuguese language models. They started with a massive collection of Portuguese documents that had already been graded for quality, ranging from low to high based on how well they were written and how useful the information was. From this collection, they created three distinct groups of data: one group containing only the highest-quality documents, another with the lowest-quality documents, and a third group that was a random mix of everything. For each of these groups, they used a separate artificial intelligence model to rewrite the text into four different styles, such as a simplified version for easier reading, a formal encyclopedia style, a technical version, and a question-and-answer format. This process generated a huge amount of new, synthetic text. They then trained two different computer models—one smaller and one larger—on these rewritten datasets to see how the combination of original quality and rewriting affected the final results.
The findings revealed a clear and surprising pattern that challenges the hope that rewriting could fix bad data. When the researchers used the larger model, the rewriting technique acted as a multiplier of quality rather than a fixer of flaws. The model trained on the high-quality documents that had been rewritten performed significantly better than the same model trained on the original, unmodified high-quality text. However, when they applied the same rewriting process to the low-quality documents, the improvement was barely noticeable. The rewritten low-quality text did not catch up to the rewritten high-quality text; instead, the gap between them remained wide. This suggests that the rewriting process is most effective when it takes already excellent material and makes it even better, rather than trying to rescue poor material. The random mix of documents fell right in the middle, showing that the benefit of rewriting grows steadily as the quality of the source text improves.
This effect, however, depended heavily on the size of the model being trained. When the researchers repeated the experiment with a much smaller model, the clear distinction between high and low quality disappeared. In this smaller setting, the rewritten high-quality text did not outperform the original low-quality text in a consistent way. The smaller model seemed unable to fully utilize the complex structural changes introduced by the rewriting process, or perhaps it simply learned better from the raw variety of the unedited data. This indicates that the power of synthetic rewriting is not a universal rule but a phenomenon that scales with the size of the artificial intelligence system. For smaller models, the extra complexity of rewritten text may not provide a benefit, while for larger, more capable models, rewriting high-quality data becomes a powerful tool for boosting performance.
The study also looked at specific types of tasks, such as answering exam questions or understanding social media posts, to see where the rewriting helped the most. The improvements were most dramatic in areas requiring structured knowledge and cultural understanding, like answering exam questions or reasoning through complex scenarios. In these areas, the model trained on rewritten high-quality data soared ahead of all others. Interestingly, for tasks involving general knowledge, the rewriting sometimes made things slightly worse, suggesting that the process can occasionally introduce noise or distort facts. For social media tasks, the original low-quality data performed surprisingly well on its own, likely because the messy, informal nature of social media text is naturally similar to the unedited web data the models were originally trained on.
Ultimately, the research demonstrates that while synthetic rewriting is a powerful technique, it is not a magic wand that can turn poor data into gold. Instead, it functions as a quality multiplier, amplifying the strengths of good data while leaving the weaknesses of bad data largely untouched. This insight is vital for developers working with languages that have fewer resources than English. It suggests that the most effective strategy is to carefully select the best available text first and then use rewriting to enhance it, rather than relying on rewriting to compensate for a lack of quality. The results also highlight that the size of the model matters; what works for a large, sophisticated system may not work for a smaller one. By clarifying how data quality and rewriting interact, this work provides a clearer path for building better artificial intelligence systems for Portuguese and, potentially, for other languages facing similar resource constraints.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.