Prostate cancer is the fifth-leading cause of cancer-related death in men. Current standard-of-care treatment schedules are based on the maximum tolerated dose (MTD) principle, in order to maximize cell kill and thereby the chance of cure. However, strategies that are based on the principle of ‘competitive release’, wherein it is hypothesized that MTD treatment removes the resource competition that restricts the growth of a drug-resistant subpopulation, have been shown to extend the time to progression (TTP) in vivo for breast cancer, ovarian cancer and melanoma. Deep reinforcement learning (DRL) models have been successfully applied to problems in healthcare ranging from automated medical diagnosis to personalized treatment regimes. However, previous work to apply conventional reinforcement learning techniques to treatment scheduling in oncology neglects the impact of treatment resistance and has not demonstrated a capability to outperform standard of care treatment techniques. We introduce a DRL model that is trained on a virtual patient to learn novel treatment strategies with minimal human direction. The virtual patient dynamics are derived from a Lotka-Volterra Model. Training on a single patient profile, the DRL model achieves an average TTP of 745 days over 100 independent iterations (95% confidence interval: (688d, 802d)). In comparison, continuous therapy and conventional adaptive therapy result in progression after 450 and 662 days on average, respectively. By applying the DRL to a virtual patient cohort with a range of characteristics, we demonstrate that it is robust to changes in tumor parameters and dynamics. We then propose a framework for developing personalized treatment schedules, which outperforms both conventional treatment strategies and non-personalized DRL models. The approaches presented here are highly generalizable to other forms of cancer, and applicable in any context where total drug load must be managed, either to prevent treatment resistance or to reduce adverse side-effects. The application of DRL models allows for treatment optimization in a wide range of clinical settings.
© 2026 - The Mathematical Oncology Blog
© 2026 - The Mathematical Oncology Blog