Welcome to the Differential-TD Convergence repository, which accompanies the paper: Almost Sure Convergence of Differential Temporal Difference Learning for Average Reward Markov Decision Processes.
This repo contains the experimental results supporting the paper on differential TD learning. It implements an off-policy
We recommend the use of uv to manage packages, and the dependencies by navigating to the directory, and running To run experiments with custom configurations, use:
uv sync --all-extrasThen run the experiments with
python differential_td.py
