Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Differential-TD Convergence

Welcome to the Differential-TD Convergence repository, which accompanies the paper: Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes.

Introduction

This repo contains the experimental results supporting the paper on differential TD learning. It implements an off-policy $n$-step differential TD agent and empirically demonstrates convergence beyond the theoretical conditions established in the paper. In particular, the experiments show stable convergence on gridworld environments across a range of step sizes $\eta$ and $n$-step parameters, and highlight remaining challenges in the off-policy setting. Namely, executing the code will produce the following figures:

Convergence with eta Convergence with n

Usage

We recommend the use of uv to manage packages, and the dependencies by navigating to the directory, and running To run experiments with custom configurations, use:

uv sync --all-extras

Then run the experiments with

python differential_td.py

About

Experiments for: Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages