Right on Time: Structured Robot Learning of Planning and Control Strategies for Signal Temporal Logic
Name
meng-mengyue-phd-aa-2025-thesis-updated.pdf
Size
16.26 MB
Format
Adobe PDF
Checksum (MD5)
b5befc78265815b1eb351dc045876521
Author(s)
Meng, Yue
Advisor(s)
Fan, Chuchu
Date Issued
February 2026
Publisher
Massachusetts Institute of Technology
Abstract
Long-horizon tasks with complex temporal constraints are ubiquitous in robotics applications, such as autonomous driving and manipulation. Signal Temporal Logic (STL) provides a formal and expressive language for specifying such tasks; however, planning executable trajectories or generating control policies to satisfy these specifications remains a significant challenge. Traditional methods are usually restricted to linear systems or simple specifications, as rigorously solving general STL tasks has been proven to be NP-hard. Furthermore, the inherent non-Markovian nature of temporal logic formulas fundamentally hinders the direct plug-and-play application of standard reinforcement learning (RL) algorithms. This thesis posits that, by explicitly leveraging the unique structure of STL specifications, we can develop learning-based methods with superior effectiveness, generalization, and versatility. Guided by this design philosophy, we introduce three distinct yet complementary approaches in this thesis, each tailored to different assumptions about the system dynamics and data availability. (1) For systems with differentiable dynamics, we propose to leverage the smoothed robustness score to train or refine the neural network predictive controller to generate rule-compliant trajectories. This training paradigm is suitable for a wide range of dynamics and temporal logic formulas. (2) When expert demonstrations for different STLs are available, we utilize the tree structure of temporal logic and develop a graph-encoded flow matching framework that can learn a universal policy for general tasks. (3) With unknown dynamics and no demonstrations, we decompose the specification into subgoals and invariant constraints governed by time variables, and design a hierarchical time-conditioned multi-goal RL framework to solve STL tasks. The efficacy of these algorithms is validated through experiments across a wide range of challenging domains, including multi-agent autonomous driving, ship tracking control, robot arm manipulation, drone navigation, and quadrupedal navigation. Our methods consistently demonstrate superior task satisfaction rates, learning efficiency, and scalability to complex, deeply nested formulas when compared to existing baselines.
MIT Department
Massachusetts Institute of Technology. Department of Aeronautics and Astronautics
Terms of Use
Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)
Copyright retained by author(s)
Persistent DSpace Link