RL + Transformer = A General-Purpose Problem Solver

Micah Rentschler; Jesse Roberts

doi:10.18653/v1/2025.realm-1.29

RL + Transformer = A General-Purpose Problem Solver

Abstract

What if artificial intelligence could not only solve problems for which it was trained but also teach itself to tackle novel tasks? In this paper, we finetune Llama 3.1 using reinforcement learning on the grid-world game Frozen Lake and investigate its ability to solve maps it has never encountered—a phenomenon recently termed In-Context Reinforcement Learning (ICRL). Without additional training, the transformer demonstrates the capacity to adapt to both in-distribution and out-of-distribution environment parameterizations. Moreover, it remains effective when trained on data that blends optimal and suboptimal behavior, combines strategies from its context (behavior-stitching), and dynamically adapts to non-stationary environments. These proof-of-concept findings suggest that in-context learning via reinforcement-tuned transformers may form the basis of a promising general-purpose problem-solver.

Anthology ID:: 2025.realm-1.29
Volume:: Proceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025)
Month:: July
Year:: 2025
Address:: Vienna, Austria
Editors:: Ehsan Kamalloo, Nicolas Gontier, Xing Han Lu, Nouha Dziri, Shikhar Murty, Alexandre Lacoste
Venues:: REALM | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 401–410
Language:
URL:: https://preview.aclanthology.org/mtsummit-25-ingestion/2025.realm-1.29/
DOI:: 10.18653/v1/2025.realm-1.29
Bibkey:
Cite (ACL):: Micah Rentschler and Jesse Roberts. 2025. RL + Transformer = A General-Purpose Problem Solver. In Proceedings of the 1st Workshop for Research on Agent Language Models (REALM 2025), pages 401–410, Vienna, Austria. Association for Computational Linguistics.
Cite (Informal):: RL + Transformer = A General-Purpose Problem Solver (Rentschler & Roberts, REALM 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/mtsummit-25-ingestion/2025.realm-1.29.pdf

PDF Cite Search Fix data