Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
This paper challenges the assumption that sophisticated prompting techniques enhance Large Language Model performance by demonstrating through a comprehensive empirical study that baseline prompting consistently outperforms complex methods, with only minimal expert or inductive framing yielding slight improvements, thereby suggesting the field should prioritize genuine model advancement over prompt engineering.