TM-RUGPULL: A Temporary Sound, Multimodal Dataset for Early Detection of RUG Pulls Across the Tokenized Ecosystem
The paper introduces TM-RugPull, a rigorously curated, multimodal dataset of 1,028 token projects designed to overcome data leakage and modality limitations in existing resources, thereby establishing a new benchmark for the early detection of rug-pull scams across the tokenized ecosystem.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine the world of cryptocurrency as a massive, bustling digital marketplace. In this market, people create new "tokens" (digital coins) to sell to investors. Sometimes, the creators are honest and build something valuable. Other times, they are scammers who run a "Rug Pull."
Think of a Rug Pull like a carnival game operator who sets up a booth, takes everyone's money, and then suddenly rips the carpet out from under the players and runs away with the cash before anyone realizes the game is rigged.
The paper you shared, TM-RugPull, is about building a better "training manual" for computers to spot these scammers before they run away. Here is the breakdown in simple terms:
The Problem: The "Backwards" Training Manual
Currently, the tools computers use to find scammers are flawed. It's like trying to teach a security guard to spot a thief by showing them pictures of the thief after they have already stolen the money and fled.
- The Flaw: Old datasets use clues that only appear after the scam is done (like the price crashing or the money disappearing).
- The Result: Computers trained on this data think they are great at predicting scams, but in the real world, they fail because they are looking at clues that don't exist until it's too late. They also mostly only look at one type of project (DeFi), missing scams in memes, games, or celebrity tokens.
The Solution: The "Time-Travel" Dataset
The authors created a new dataset called TM-RugPull. They designed it like a strict time-travel rule: "You can only look at information available before the disaster happens."
Here is how they built it:
1. A Diverse Crowd (The 1,000 Projects)
Instead of just looking at serious financial apps, they gathered 1,000 different projects. Imagine a lineup that includes:
- Serious financial apps (DeFi).
- Funny internet jokes (Meme tokens).
- Digital art and games (NFTs).
- Celebrity-branded coins.
This ensures the computer learns to spot scammers in any type of digital project, not just the serious ones.
2. The "Midpoint" Rule (Stopping Time Leakage)
To make sure the computer doesn't cheat by looking at the future, the authors picked a specific "Midpoint" in time for every project.
- The Rule: The computer is only allowed to see data from the first half of the project's life (before the midpoint).
- The Check: If the project turns out to be a scam later, the computer is only graded on whether it could have predicted it using only the early clues. If the computer used a clue that appeared after the scam started, it's disqualified. This is called preventing "temporal leakage."
3. Three Eyes on the Market (Multimodal Data)
To get a full picture, the dataset looks at the projects through three different "lenses" simultaneously:
- Lens 1: On-Chain Behavior: Watching the actual movement of money and tokens on the blockchain (like watching the cash register).
- Lens 2: Smart Contract Metadata: Reading the "rules" written in the code itself (like reading the fine print of a contract).
- Lens 3: OSINT (Open Source Intelligence): Checking social media, Google searches, and news (like listening to the gossip in the town square).
- Why it matters: Often, people start talking about a scam on Twitter or Google before the money actually disappears on the blockchain. This dataset catches those early whispers.
4. The Human Detective Team (Manual Labeling)
The authors didn't just let a computer decide who is a scammer. They used a team of human experts.
- The Process: They watched projects for 72 hours after things went wrong. If the liquidity vanished, activity stopped, and the price froze, they labeled it a "Rug Pull."
- The Safety Net: They double-checked this against known scam databases and had two experts agree on the verdict to ensure no mistakes were made.
The Result
The paper presents a clean, honest, and diverse training set for AI.
- It has 1,000 projects (about 60% scammers, 40% honest).
- It covers 5 different blockchains (like Ethereum and Binance).
- It strictly forbids "cheating" by using future information.
In short: The authors built a "flight simulator" for AI security guards. Instead of training them on videos of crashes that already happened, they trained them on the pre-crash signals so the guards can actually stop the plane from crashing in the first place. They have made this simulator and the code to build it available for everyone to use.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.