Energy Grids

Learning Data-Center Flexibility for Grid Demand Response with RMABs

A critical note on modeling data centers as restless bandit arms for grid demand response, and on where the finite-state RMAB abstraction is useful or fragile.

Problem: data centers as flexible but private grid resources

The paper studies a grid-demand-response setting in which data centers can reduce short-term electricity consumption by delaying or reshuffling virtual-machine jobs. From the grid operator’s perspective, this is attractive: data centers are large, controllable loads, and even a modest temporary reduction can be valuable during a peak or congestion event.

The hard part is that the grid operator does not see the internal scheduling state of each data center. It does not know the full job queue, delay sensitivity, QoS cost, thermal condition, or server-level availability. The operator only sees an aggregate or batch-level context and must decide which data centers should receive a flexibility request in each round.

The simplified decision problem is:

maxπ Eπ [ ∑t=0∞ βt ∑i=1N ri,t ai,t ] subject to ∑i=1N ai,t ≤ Nt , ai,t ∈ {0,1} .

Here each data center is an arm. Activating arm i means sending a flexibility request to that data center. The reward is the electricity-saving benefit minus the QoS or delay penalty caused by the rescheduling action.

This is a natural restless multi-armed bandit formulation because inactive data centers do not freeze. Their internal queues keep evolving even when the grid operator does not request flexibility.

State and reward abstraction

The paper makes the data-center state finite by representing it as the current position in a circular VM-job queue. A state is not the full physical or operational state of the data center; it is the current batch position:

s = 1: [job 1, job 2, ..., job Nj]
s = 2: [job 2, job 3, ..., job Nj+1]
...

When the operator activates a data center, the data center can choose lower-power jobs from a lookahead window and thereby reduce instantaneous power. The reward is roughly:

r(s) = max { 0 , LMP · ( Pdef(s) - Psel(s) ) - Cdelay(s) } .

The important modeling choice is that the reward and transition dynamics are attached to this finite queue state. This makes the RMAB solvable and learnable, but it also compresses a very rich operational system into a small discrete representation.

Whittle index view

Solving the full dynamic program over all data centers is intractable because the joint state space grows exponentially with the number of arms. The Whittle-index relaxation turns the global allocation problem into many single-arm subsidy problems.

The intuition is:

Original problem:
  choose the best subset of data centers jointly

Whittle relaxation:
  ask each data center-state independently:
  how much passive subsidy would make passivity as attractive as activation?

For one arm, the active and passive values under a passive subsidy λ can be written schematically as:

Q1(s) = R(s,1) + β ∑s′ P1(s,s′) V(s′) , Q0(s) = R(s,0) + λ + β ∑s′ P0(s,s′) V(s′) .

The Whittle index is the smallest passive subsidy that makes passive action preferable. A large index means that the operator would need to pay a large subsidy to justify not activating that data center-state; equivalently, activation is valuable.

Learning: Thompson-Whittle with a trust-mixed exploration stage

Classical Whittle-index policies assume that rewards and transition probabilities are known. In the data-center setting, the operator does not know them. The paper therefore uses a Thompson-sampling style model-learning approach.

For each arm, the unknown model is treated as a stationary parameter:

ri,t = Ri ( si,t , ai,t ) + εi,t , si,t+1 ∼ Pi ( · | si,t , ai,t ) .

The transition model uses Dirichlet-Categorical uncertainty, and the reward model uses Gaussian uncertainty. Each round samples a plausible model from the posterior, computes Whittle indices under that sampled model, and activates the top-ranked arms.

Pure Thompson-Whittle has a clear failure mode. Early in learning, many state-action transitions have barely been observed. The sampled transition model can be unstable, and the resulting Whittle-index ranking can be wrong. The paper’s proposed Trust-Mixed Thompson-Whittle approach adds a decaying UCB-based greedy score:

early rounds:   global/local UCB dominates
middle rounds:  state-level UCB and Whittle both matter
late rounds:    Whittle index dominates

This is best read as a practical mixed strategy: short-term robust exploration first, transition-aware RMAB control later.

Experimental reading

The experiments use Microsoft Azure VM traces. The paper filters low-usage or low-utilization jobs and estimates power consumption and QoS cost, with GPU power modeled using hardware specifications such as NVIDIA A100 static and maximum power.

The main comparisons are Oracle Whittle, pure Thompson-Whittle, the proposed trust-mixed method, a state-only Thompson-style baseline, and EXP4-style expert mixing. The strongest empirical message is not that the proposed method changes the long-run RMAB theory. It is that pure Thompson-Whittle can be fragile in the early learning stage, and a scheduled UCB mixture can reduce that transient regret.

The reported improvement over pure Thompson-Whittle is positive but not overwhelming in every table. The supplied summary notes a roughly 0.17-2.14% improvement over TW in some settings. That should be interpreted as an early-learning robustness gain, not as a broad theoretical dominance result.

The noisy-context experiment is also subtle. When observation noise is high, the proposed method can outperform an “Oracle Whittle” baseline. This does not mean it beats an oracle with access to the true hidden state. It means the oracle baseline still performs a state-indexed Whittle lookup using corrupted observed state, while the UCB mixture can exploit aggregated statistics that are less brittle under noisy state labels.

Where the abstraction is useful

The application idea is clean: if a grid operator cannot directly observe or query all internal data-center costs, it may learn which data centers tend to be useful in which coarse states. The RMAB formulation also captures an important feature that a myopic contextual bandit misses: activation affects future state distributions, not only immediate reward.

This makes the paper a useful computational proof-of-concept for grid-data-center coordination under limited observability. It gives a plausible algorithmic template:

aggregate batch context
  -> posterior over reward and transition models
  -> sampled single-arm MDPs
  -> Whittle-index ranking
  -> activate a limited number of data centers
  -> observe feedback and update

For an application review, the most valuable contribution is this mapping from privacy-constrained data-center flexibility to a learning RMAB structure.

Limitations: state abstraction, reward stationarity, and incentives

The weakest part of the application story is the gap between the real data-center state and the finite circular-queue state. A real operational state would include CPU and GPU utilization, memory pressure, VM queue composition, job deadlines, QoS sensitivity, thermal state, server availability, network constraints, electricity price, and perhaps carbon or local congestion signals. That is continuous, hybrid, high-dimensional, and partly private.

The paper’s finite state is therefore a coarse abstraction:

si,treal ↦ s^i,t = ψ (xi,t) .

This abstraction is valid only if the current batch position is sufficient to explain reward and transition behavior. Nonstationary workload arrivals, hidden QoS sensitivity, changing electricity prices, thermal constraints, or regime shifts can break that Markov abstraction.

The same issue appears in the reward model. Thompson sampling is not a magic solution for a reward that changes arbitrarily over time. In the paper, Thompson sampling learns a stationary unknown reward and transition model. If the true reward is instead driven by an external regime variable,

ri,t = Ri ( si,t , ai,t ; ηt ) + εi,t ,

then electricity price, workload regime, SLA condition, carbon intensity, or local grid congestion must either be included in the state/context or treated as stationary noise. Otherwise the posterior simply learns an averaged stationary model over a misspecified process.

There is also a procurement question. If each data center can truthfully report real-time flexibility, cost, and reliability, the grid operator could solve a direct flexibility procurement problem instead of learning an RMAB policy:

mina ∑i ci,t ai,t subject to ∑i Δ Pi,t ai,t ≥ DtDR , ai,t ∈ {0,1} .

That market-based formulation may be more natural when truthful real-time bids are available. The RMAB approach becomes more compelling when bids are unavailable, delayed, strategically filtered, or too privacy-sensitive to expose directly.

Finally, reward observability is under-specified. If the operator does not see internal QoS cost, how does it observe the net reward “power saving minus delay penalty”? If the data center reports the QoS penalty, incentive compatibility becomes central. If it does not, the operator may only learn grid-side power response, not true social reward. This creates tension with the privacy-preserving claim.

Takeaway

This paper is best read as an application-oriented RMAB learning proof-of-concept, not yet as a complete grid-data-center coordination architecture. The finite circular-queue model makes the problem tractable and allows Thompson-Whittle learning, while the trust-mixed UCB stage addresses a real early-learning instability.

The application claim would be stronger with three additional pieces: a richer state or context model for continuous data-center operation, a clear settlement mechanism for observing net reward without violating privacy, and an incentive-compatible reason why the grid operator should learn activations instead of receiving real-time flexibility bids.

References

Ding, Y., Chen, Z., & Magnanti, T. (2026). Robust Restless Multi-Armed Bandit for Data Center Flexibility Services Through Virtual Machine Scheduling. arXiv preprint arXiv:2605.19116.