User eXperience Perception Insights Dataset (UXPID): Synthetic User Feedback from Public Industrial Forums
This paper introduces UXPID, a synthetic dataset of 7,130 anonymized user feedback branches from industrial forums, annotated by LLMs to support research in UX analysis, requirements extraction, and AI-driven feedback processing while addressing privacy and licensing constraints.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Imagine you are the captain of a massive ship, but instead of sailing the ocean, you are navigating the complex world of industrial software and hardware. Your crew (the users) constantly sends you notes, complaints, and ideas through a giant, chaotic bulletin board. Some notes are scribbled on napkins, some are shouted over the noise, and many are written in a secret code only the writers understand.
This paper introduces a new tool called UXPID (User eXperience Perception Insights Dataset) to help captains like you make sense of that chaos.
Here is the story of how they built it, what's inside, and why it matters, explained simply:
1. The Problem: The "Noise" in the Room
For years, companies have had a goldmine of information: thousands of user comments on public forums. But this gold is buried under dirt.
- The Mess: The comments are unstructured, messy, and scattered.
- The Privacy Wall: You can't just grab the real notes because they contain private names, company secrets, and specific product codes. It's like trying to read a diary that has been locked in a safe.
- The Missing Map: Existing tools for analyzing these notes are like trying to find a needle in a haystack without a magnet. They often miss the severity of a problem (is it a tiny scratch or a hole in the hull?) or the specific topic (is this about the engine or the navigation system?).
2. The Solution: The "Synthetic Mirror"
The researchers created UXPID, which is like a perfectly polished, anonymized mirror of that messy bulletin board.
- How they made it: They took 7,130 real conversation threads from an industrial automation forum.
- The Magic Trick (AI): They used a super-smart AI (a Large Language Model) to act as a "translator" and "privacy guard."
- Step 1: The AI read the real, messy comments and extracted the meaning: "The user is angry," "The product is broken," "They need a new feature."
- Step 2: The AI then rewrote the comments. It kept the meaning exactly the same but swapped out all the secrets. Instead of "John Smith from Siemens," it wrote "[User Name] from [Company Name]." Instead of "Product X version 2.1," it wrote "[Product Name] [Version No]."
- The Result: You get a dataset that feels exactly like real life but is safe to share with anyone. It's like serving a delicious meal where the ingredients are real, but the chef has removed the allergens.
3. What's Inside the Box?
The dataset isn't just a pile of text; it's a structured library. Every single conversation thread comes with a "tagged summary" that includes:
- The "Pain Points": What is the user frustrated about?
- The "Gain Points": What do they love?
- The "Severity": Is this a minor annoyance or a critical emergency?
- The "Sentiment": Are they happy, sad, or neutral?
- The "Topics": What category does this fall under?
Think of it as taking a chaotic shouting match and turning it into a neatly organized filing cabinet where every complaint is labeled with a color-coded tag.
4. Did It Work? (The Test Drive)
The researchers didn't just build the box; they tested if it actually helps computers learn. They taught two different types of "students" (computer models) using this new dataset:
- The Old Student: A traditional method that looks for keywords (like counting how many times "broken" appears).
- The New Student: A modern AI (DistilBERT) that understands context and nuance.
The Results:
- The New Student learned much faster and better. It could correctly guess the topic of a conversation and the user's mood with high accuracy.
- The Old Student struggled, often guessing wrong or getting confused by rare words.
- Key Finding: The new AI was so good at understanding the meaning behind the words that it could predict what the user wanted even when the words were slightly different.
5. The "Fidelity" Check: Did We Lose the Soul?
A major concern was: "When we changed the names and numbers, did we change the feeling of the text?"
- They compared the original messy notes with the clean, synthetic ones.
- The Verdict: The structure of the conversation stayed exactly the same. The number of questions asked and the length of the replies were identical.
- One Small Change: The synthetic text became slightly more "formal." The AI removed some of the shouting (exclamation marks) and all-caps yelling to protect privacy. This means if you train a computer on this, it might be slightly less sensitive to "shouting" in real life, but the core message remains intact.
Summary
In short, this paper presents a safe, high-quality training manual for computers. It allows researchers and companies to teach AI how to listen to customer feedback, understand how angry or happy they are, and figure out what they need—without ever seeing a single real person's private name or secret company data. It turns a chaotic crowd of voices into a clear, actionable signal.
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.