Tacrolimus RL Dosing Agent (LSTM-PPO)

LSTM-augmented PPO agent for adaptive tacrolimus dose adjustment under CYP3A4-mediated drug-drug interactions (DDIs).

Trained on a two-compartment PBPK simulator spanning 8 CYP3A4 inhibitors (fluconazole, voriconazole, posaconazole, ketoconazole, verapamil, diltiazem, erythromycin, clarithromycin) with full patient variability (CYP3A5 genotype, hematocrit recovery, steroid taper, non-adherence, food effects).

Observation space: 8-dim POMDP (trough conc, DDI flag, episode day, dose history) Action space: 9 discrete dose levels (0–5 mg BID) Episode length: 90 days Training: 2M timesteps, 8 parallel envs

Paper: Reinforcement Learning for Adaptive Tacrolimus Dosing with Multi-Drug Interaction Management (ICML 2026)

Downloads last month
484
Video Preview
loading