Online Reinforcement Learning for Digital Health Interventions in the Dyadic Setting
We present our ongoing work on the development of an online reinforcement learning (RL) algorithm that personalizes delivery of interventions sequentially, over time, to members of a dyad. The RL algorithm is one component in the design of the digital health intervention. Different RL components target different elements of the dyad; these elements are a target individual, a care partner and their relationship. The RL algorithm is a multi-agent RL algorithm in which the 3 agents make decisions on the 3 elements of the dyad. We incorporate domain knowledge in the form of approximal causal directed acyclic graphs to speed up online learning in this sparse data setting. This work is motivated by our development of the ADAPTS-HCT multi-agent RL algorithm, designed to improve medication adherence by young adults who have undergone a blood and bone marrow transplant. The RL algorithm will be deployed in a clinical trial in late summer 2026.