LLM & Probabilistic Approaches

Tolerance Ball Acquisition for Specification-Driven Inverse Design

A critical note on Tolerance Ball acquisition: a clean probability-of-feasibility objective for specification-driven inverse design, but not a direct optimizer of diversity, boundary coverage, or global feasible-set recovery.

The problem is not the usual Bayesian optimization problem of finding one best design. In many materials and process-design settings, the useful question is closer to this:

Find many designs x such that f(x) lies inside a target property window.

The target may be a glass-transition-temperature interval, a bandgap range, a viscosity window, or a molecular-weight-distribution profile that only needs to be close enough. Exact matching is rarely necessary. Once several candidates satisfy the specification, the real choice often depends on cost, synthesis difficulty, robustness, toxicity, scale-up risk, or operational convenience.

That makes the objective different from ordinary extremum search. Standard EI, PI, and UCB are designed to locate a maximum or minimum. Multi-objective Bayesian optimization usually searches for a Pareto frontier. But a specification window can contain many useful points that are Pareto-dominated and still practically valuable. The paper behind this note reframes the task as specification-driven inverse design: instead of asking for the best point, ask for a set of valid points.

Why Standard BO Is Misaligned

In standard Bayesian optimization, the next design is chosen by maximizing an acquisition function over the current posterior. Expected Improvement, for example, rewards points that may improve on the current incumbent. That is sensible when the objective is a single best value.

For target-window discovery, this incentive can be wrong. If one feasible point has already been found, EI may keep searching nearby for a slightly better discrepancy rather than looking for other disconnected feasible regions. The method becomes good at local refinement, but not necessarily good at collecting many practically different valid designs.

Constrained BO is closer. Constrained EI often multiplies an improvement term by a probability of feasibility. But in that setting feasibility is still a side condition attached to an optimization objective. The Tolerance Ball idea removes the separate objective and makes the probability of satisfying the target window the acquisition itself.

So the acquisition is best read as a target-centered probability of feasibility.

That is a useful shift. It is also narrower than the phrase “discovering diverse designs” might suggest.

The Tolerance Ball Objective

Let x be a design variable and f(x) be a K-dimensional property vector. A target vector y_tgt is given, and a design is valid if the predicted property lies inside a ball of radius epsilon around the target:

||f(x)-ytgt|| 2 ≤ ε2.

With multiple targets, each target defines its own valid set. The paper uses a shared Gaussian-process posterior for the property map and then computes a target-specific acquisition for each target.

The architecture is simple. Independent GPs are fitted for the output dimensions. Their posterior means and variances are combined into an approximate multi-output Gaussian. The acquisition at a candidate x is the posterior probability that f(x) falls inside the tolerance ball.

Under the isotropic approximation

f(x)|D ≈ N( μ(x), η2(x)I ),

where eta squared is the average of the output-wise posterior variances, the tolerance-ball probability can be computed with a noncentral chi-square CDF. This is the clean mathematical part of the method. It avoids posterior sampling and directly scores the next point by its probability of being valid.

What TB Actually Optimizes

The acquisition is exactly aligned with a one-step utility:

u(x) = 1 { ||f(x)-ytgt|| 2 ≤ ε2 }.

The expected value of this utility is the probability that the next evaluated point is valid. Maximizing the TB acquisition therefore maximizes the posterior probability of getting a valid hit in the next experiment.

That is a precise and defensible objective. If the goal is to harvest valid candidates quickly near a known target region, TB is well matched to the task.

But it does not optimize a long-horizon portfolio objective. It does not directly maximize coverage of the feasible set. It does not directly maximize chemical diversity, process diversity, or the number of disconnected feasible components discovered. Those goals would require terms that depend on previously found valid designs, distances between candidates, feasible-region geometry, or an explicit exploration policy.

This is the central distinction: TB is a valid-hit acquisition, not a diversity acquisition.

Where Exploration Comes From

The paper discusses exploration-exploitation behavior, but TB does not reward uncertainty in the usual optimistic sense. In a one-dimensional target interval [a,b], the acquisition is the posterior probability mass inside the interval. If the posterior mean is already inside [a,b], reducing variance usually increases that probability. Large uncertainty spreads probability mass outside the target window.

Uncertainty helps only in a particular case: when the posterior mean is near the target but outside it, some uncertainty can push probability mass into the valid region. That is not global exploration. It is target-adjacent uncertainty exploitation.

This matters because inverse design often contains disconnected feasible islands. Suppose the valid set has several components, but the initial data are concentrated near one component. TB can keep selecting points around the known component because that region has the largest posterior probability of validity. Preventing exact duplicate measurements is not enough. In a continuous space, many near-duplicate candidates can still live in the same local basin.

So the method may collect many valid points without discovering the full structure of the valid set.

Diversity Is Evaluated, Not Optimized

The most important limitation is that diversity appears mainly as an evaluation metric. In continuous domains, the paper uses a distance-based uniqueness notion. In discrete libraries, it counts distinct valid candidates. These are reasonable reporting metrics, but the acquisition itself does not include a diversity penalty or reward.

That creates an objective-metric mismatch. The method is evaluated by asking whether it found many separated valid designs, but the actual decision rule asks which next point has the highest posterior probability of being valid.

A direct diversity-seeking acquisition would need a term such as distance from previously found valid candidates, coverage of underexplored feasible regions, novelty in a molecular graph kernel, scaffold diversity, synthesis-route diversity, or another task-specific notion of difference. Without such a term, any observed diversity is indirect. It may come from posterior geometry, target multiplicity, finite-library structure, or the duplicate-avoidance rule.

This does not make TB useless. It just narrows the claim.

Boundary Learning Is a Different Problem

TB also should not be confused with level-set estimation or contour learning. If the valid set is defined by a tolerance ball, the boundary is where the distance from the target equals epsilon. Near that boundary, the probability of validity can be around one half. In the interior, it can be close to one.

TB prefers the high-confidence interior. A boundary-learning method would deliberately sample near uncertain boundaries to learn the shape, volume, and extent of the feasible set. TB is closer to high-confidence valid-point harvesting.

This distinction is important in safety-critical or regulation-constrained settings. If one needs to map the boundary of the safe operating envelope, TB is not the right objective by itself. If one only needs additional candidates that are likely to satisfy a known specification, TB can be appropriate.

High-Dimensional Outputs

The high-dimensional output case is delicate. In the molecular-weight-distribution example, the output can be a 100-dimensional compositional vector. Treating this as an ordinary Euclidean vector with independent output GPs and an isotropic covariance approximation is strong.

There are two problems.

First, tolerance-ball probability in high dimension can collapse rapidly as posterior uncertainty increases. Even when the posterior mean is exactly at the target, the probability mass inside a small ball shrinks sharply with dimension. In a 100-dimensional output space, TB can become strongly biased against uncertain regions.

Second, a molecular-weight distribution is compositional. Its bins sum to one, and the bins are correlated: increasing one bin usually forces decreases elsewhere. Independent output GPs and an isotropic covariance approximation ignore this structure. The closed-form acquisition is computationally convenient, but the approximation may not respect the geometry of the output.

This is not a small modeling detail. In high-dimensional correlated outputs, the acquisition can be dominated by the covariance approximation rather than by the real design question.

What Is Valuable

The contribution is not that TB creates a completely new principle for Bayesian optimization. It is the combination that is useful:

specification-driven inverse design, a tolerance-region utility, a sampling-free multi-output acquisition, shared posterior learning across targets, and materials-oriented case studies.

The strongest practical idea is to stop pretending that the best point is always the right target. In many design problems, the engineer wants several candidates that satisfy a specification, then chooses among them using external criteria. TB expresses that immediate objective more honestly than an extremum-seeking acquisition.

The closed-form acquisition is also attractive. Methods that estimate feasible sets by posterior sampling can be expensive. A noncentral chi-square CDF is much cleaner, provided the Gaussian and isotropic assumptions are acceptable.

Reading the Experiments

The reported benchmarks are broad: synthetic functions, pool-based materials datasets, polymerization design with molecular-weight-distribution targets, and sequence-defined oligomer libraries. The compared methods include TB, HV, EI, LCB, BAX, and random sampling. In the reported results, TB obtains the best average rank overall and strong diversity scores in the polymerization case.

That is meaningful evidence that the acquisition works well in the tested settings. The polymerization result is particularly useful from a process-design viewpoint: finding different reaction conditions that lead to similar target MWDs can give a real operator more flexibility.

But the benchmark design should be read carefully. In pool-based datasets, targets selected from the observed output distribution are known to be feasible after the fact. Real inverse design often starts from an external requirement: a market target, a device-level specification, or a regulatory threshold. Such targets may be far from the training distribution, sparsely feasible, or infeasible.

The oligomer case also shows a weaker regime for TB. When a finite library already contains many valid candidates, random sampling can be competitive for some targets and the method gap shrinks. That does not refute TB. It shows that its value depends on scarcity, model calibration, and the geometry of the candidate set.

Assessment

The clean guarantee is narrow. If the posterior approximation is correct, TB computes the probability that a candidate lies inside the tolerance ball. And because that probability is the expected value of a valid-hit indicator, TB is exactly aligned with one-step valid-hit maximization.

What is not guaranteed is just as important: global feasible-set recovery, boundary coverage, disconnected-component discovery, batch diversity optimality, long-horizon portfolio diversity, regret bounds, or global convergence.

So the right reading is modest. TB is a good acquisition when the feasible region is at least partly known, the model is reasonably calibrated, and the goal is to collect valid candidates efficiently near target specifications. It is less convincing when the design space is unknown, feasible islands are sparse and disconnected, outputs are high-dimensional and correlated, or the actual goal is useful diversity rather than valid-hit rate.

References

Jiang, S., Wu, J., Schroeder, C. M., & Webb, M. A. (2026). Range-aware Bayesian optimization for discovering diverse designs within target property windows (arXiv:2606.11574v1). arXiv.