Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
This paper introduces Policy Iteration with Human Feedback (PIHF), a framework that leverages a language-model critic and clinical expert oversight to iteratively refine natural-language policies for rare-disease diagnosis, demonstrating significant performance gains across diverse model sizes without altering their underlying weights.
Original paper licensed under CC BY 4.0 (http://creativecommons.org/licenses/by/4.0/). This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer
Drowning in papers in your field?
Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.