Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Snabbare konvergens för algoritmer för förstärkt inlärming med hjälp av förtränade stora språkmodeller som handledare med återanvändning av råd
2025 (Swedish)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesisAlternative title
Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing (English)
Abstract [en]

Reinforcement Learning (RL) has shown remarkable capabilities in solving complex decision-making problems, yet it often suffers from slow convergence and high computational demands. This study investigates the potential of using pre-trained Large Language Models (LLMs) as external tutors to accelerate RL convergence. A novel student-teacher architecture is proposed, where RL agents receive structured guidance from LLMs, including an advice reusing mechanism that stores and re-applies previously suggested actions. The effectiveness of this approach is evaluated across 49 experimental configurations, incorporating three RL algorithms (DQN, PPO, A2C), three environments (Blackjack, Connect Four, Snake), and three open-source LLMs (LLaMA3.1, Vicuna, DeepSeek-R1). Results demonstrate that LLM tutoring accelerates convergence without degrading the agents' final performance, with further improvements observed when advice reuse is employed. DeepSeek-R1, the largest tested model, achieved the most significant impact. These findings suggest a promising pathway for leveraging LLMs in RL training, highlighting opportunities for more sample-efficient, scalable, and explainable learning frameworks.

Place, publisher, year, edition, pages
2025.
National Category
Artificial Intelligence
Identifiers
URN: urn:nbn:se:hj:diva-68761OAI: oai:DiVA.org:hj-68761DiVA, id: diva2:1972566
Available from: 2025-08-04 Created: 2025-06-18 Last updated: 2025-10-13Bibliographically approved

Open Access in DiVA

fulltext(488 kB)493 downloads
File information
File name FULLTEXT01.pdfFile size 488 kBChecksum SHA-512
cf1bb473f22a1fc92a47f3b4117d1690e7076266e722689416910f39561cbfbf6c671da5f3e4c04ad2f479a38d5d87b9d94ddb0632e0661010016049fc6789fe
Type fulltextMimetype application/pdf

Artificial Intelligence

Search outside of DiVA

GoogleGoogle Scholar
Total: 496 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 790 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf