CMDP, safe RL, constrained policy optimization, Lyapunov methods, feasibility, and safety guarantees.CMDP, 안전 강화학습, 제약 정책 최적화, Lyapunov 방법, 실행 가능성, 안전 보장.
CAffNet embeds input-dependent affine constraints in a neural output layer. It guarantees feasibility for known nonempty constraint sets, but its combinatorial active-set cost and model-dependent safety assumptions remain central limitations.CAffNet은 입력 의존 affine 제약을 신경망 출력층에 삽입한다. 알려진 비공집합 제약집합에 대해서는 feasibility를 보장하지만, 조합적 active-set 비용과 모델 정확성에 의존하는 안전 가정은 여전히 핵심 한계다.
CAffNet: Hard Constraint-Affine Neural NetworksCAffNet: Hard Constraint-Affine Neural Networks
How online MDP estimation, interval uncertainty, and probabilistic shielding form a coupled loop in which blocking risky actions can also block the evidence needed to learn safety.온라인 MDP 추정, 구간 불확실성, 확률적 실딩이 결합될 때 위험한 행동의 차단이 안전을 학습하는 데 필요한 증거까지 막는 교착을 분석한다.
A critical note on enforcing input-dependent linear equality and inequality constraints in neural network outputs using a robustly feasible decision-rule anchor and minimal interpolation.강건하게 실행가능한 결정규칙 앵커와 최소 보간을 이용해 신경망 출력의 입력 의존 선형 등식 및 부등식 제약을 강제하는 방법에 대한 비판적 연구 노트.
Enforcing hard linear constraints in deep learning models with decision rulesEnforcing hard linear constraints in deep learning models with decision rules
This note reads Lyapunov-based safe policy optimization as a practical projection bridge from finite CMDP safe policy iteration to continuous-action deep reinforcement learning, while separating exact CMDP guarantees from local approximation behavior.이 노트는 Lyapunov 기반 안전 정책 최적화를 유한 CMDP의 안전 정책 반복과 연속 행동 심층 강화학습을 잇는 실용적 투영 구조로 읽는다. 동시에 정확한 CMDP 보장과 국소 근사에서의 경험적 안전성을 구분한다.
Lyapunov-based Safe Policy Optimization for Continuous Control연속 제어를 위한 Lyapunov 기반 안전 정책 최적화
Lyapunov constraints turn an expected cumulative safety budget in a CMDP into local restrictions on policy improvement. This note examines the exact certificate logic and the weaker status of neural approximations.Lyapunov 제약은 CMDP의 기대 누적 안전 예산을 정책 개선의 국소 제약으로 바꾼다. 이 노트는 정확한 인증 논리와 신경망 근사에서 약해지는 보장 수준을 구분해 읽는다.
A lyapunov-based approach to safe reinforcement learning안전 강화학습을 위한 Lyapunov 기반 접근