← Latest papers
💻 computer science

BuyTheBy: A dataset of 18,710 text-based paper mill advertisements with 51,812 timestamped prices

This paper introduces BuyTheBy, a large-scale dataset comprising over 18,000 timestamped text-based advertisements and more than 51,000 price points from seven international paper mills, designed to facilitate quantitative research into the market dynamics of academic fraud services.

Original authors: Reese AK Richardson, Spencer S Hong, Anna Abalkina

Published 2026-04-28
📖 4 min read☕ Coffee break read

Original authors: Reese AK Richardson, Spencer S Hong, Anna Abalkina

Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

Imagine the world of academic publishing as a giant, high-stakes marketplace. Usually, people pay for knowledge by doing the hard work of research, writing, and peer review. But there's a shadowy side market operating in the background: "paper mills." These are like illegal factories that sell pre-written scientific articles and, more importantly, sell spots on the author list so people can pretend they wrote them.

For a long time, trying to study how these "factories" operate has been like trying to map a black market while wearing blindfolds. Researchers knew the shops existed, but they didn't have a clear price list. Was a fake author spot $50 or $5,000? Did the price change depending on the country or the type of degree? Until now, we only had scattered rumors and isolated news reports.

Enter "BuyTheBy": The Big Price Tag

This paper introduces a new tool called BuyTheBy. Think of it as a massive, organized spreadsheet that finally lists the "menu prices" for academic fraud. The researchers collected 18,710 advertisements from seven different "paper mill" businesses operating out of seven different countries (including India, Russia, Ukraine, and Kazakhstan).

They didn't just guess the prices; they scraped the actual text from these businesses' websites and Telegram chat channels. The result is a dataset containing over 51,000 specific price points attached to specific products, like "First Author on a Medical Paper" or "Editor on a Textbook."

What's on the Menu?

The researchers found that these "factories" sell a variety of goods, though most of them focus on selling authorship spots on scientific articles.

  • The Variety: Some mills sell everything from textbook chapters to patent registrations. One business (B1) even sold spots on "India utility patents" and "UK design patents."
  • The Targets: Some mills target specific prestigious publishers (like Elsevier or Springer), while others target smaller, regional journals or conference proceedings.
  • The Currency: Prices were originally listed in local currencies (like Indian Rupees or Russian Rubles) but were converted to US Dollars to make them comparable.

What the Data Tells Us (The "Receipts")

By looking at this massive list of prices, the authors were able to draw some clear pictures:

  1. Prices are a Rollercoaster: The cost isn't fixed. Just like gas prices, the price for a fake author spot can drop or rise quickly. For example, one business dropped the price of a textbook author spot from about $180 to $84 in just six months. Another saw prices jump twice in a single month.
  2. Location Matters: A fake author spot in one country might cost half as much as the same spot in another. The researchers suggest this likely depends on how much money the customers in that specific region have.
  3. The "First Author" Premium: If you want the top spot (First Author), it costs the most. Across all the data, the median price for a First Author spot on an academic article is about $788. However, the price drops the further down the author list you go (2nd author, 3rd author, etc.).
  4. The Range is Wild: While the average is around $800, the prices are all over the place. The cheapest spot found was about $20, while the most expensive one listed was over $5,600.

Why This Matters

The paper argues that having this "price list" is a game-changer. Before, studying these criminal markets was like trying to understand an economy without knowing the value of money. Now, researchers can ask serious questions:

  • How does the price change if the journal is more famous?
  • Do prices go up when universities start demanding more publications for tenure?
  • Can publishers use this data to spot which of their journals are being targeted by these mills?

The Catch (Limitations)

The authors are honest about what this dataset can't do.

  • It only looks at text ads, not pictures or videos.
  • They didn't pick every single paper mill in the world; they picked the ones they could easily find and download.
  • They didn't track which of these fake articles actually got published (because paper mills often change the article titles to hide their tracks).

In a Nutshell

This paper is like a detective finally getting access to the ledger of a criminal organization. It doesn't stop the crime, but it gives the world a clear, data-driven look at how much these "academic fraud factories" charge, how their prices fluctuate, and how they operate across different borders. It turns a shadowy mystery into something that can be measured, analyzed, and understood.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →