Explore

See what is being averaged.

A simple nearest-neighbor prediction becomes an exact weighted estimator when averaged over subsamples.

Draw, predict, average

The horizontal coordinate is a feature; the vertical coordinate is an observed response. Choose a query and subsample size. Draw a subsample, use its nearest observation's response, and repeat.

Feature x
Exact BNNSubsample meanGenerating function

Move the point at which you predict.

Small s averages broadly; large s is more local.

Exact BNN estimate0.000
Mean of sampled predictionsNo draws yet
Subsamples drawn0
Exact rank weightsExpected selected rank

Synthetic sample, seed 2026, n = 24. The curve is μ(x) = 0.75 sin(2πx) + 0.4x with bounded generated noise. The sampled mean approximates the average; the rank-weight calculation gives it exactly.

Each draw selects a different possible neighborhood. An observation contributes whenever it is the closest member of the selected subsample. Those selection frequencies determine the estimator's weights.

Why does rank i get this weight?

Order all n observations by distance to the query. The observation at rank i is nearest in a subsample precisely when it is included and every closer observation is excluded. The other s−1 members must come from the n−i observations farther away.

wi(n,s) = C(n−i,s−1) / C(n,s)
μ̂s(x) = Σi=1n wi(n,s) Y(i).

The numerator counts subsamples in which this observation is nearest. The denominator counts all size-s subsamples. Thus the weight is a probability, and the estimate is the expectation of the nearest response under this subsampling scheme.

The calculation is exact

The weights are nonnegative and sum to one. For s > 1, ranks beyond n−s+1 have zero weight because too few farther observations remain to complete the subsample. There is no need to simulate or enumerate subsets in the implementation.

Sources: Steele (2009); JASA Section 2, equations (5)–(6). Prediction algorithm.

A larger subsample means a closer neighbor

At s = 1, each observation is equally likely to be selected, so BNN is the sample mean. At s = n, every subsample is the entire dataset, so only the closest response contributes. Between these extremes, weights decline with distance rank.

The expected rank of the selected neighbor is (n+1)/(s+1). This identity makes locality concrete: increasing s shifts probability toward the earliest ranks. Locality is also affected by the geometry and density of the features, which the rank alone does not describe.

A more local estimate can adapt to the regression function more closely while averaging fewer influential responses. The statistical bias–variance analysis depends on smoothness, the design, and how s grows with n. Read the bias and scale analysis.

Combine scales to cancel leading bias

Under the smoothness conditions in the inference paper, the leading bias has the same coefficient at different scales and varies as s−2/d. Choose two scales and coefficients that preserve constants while canceling that term.

Combine two base estimates with signed coefficients.
Two-scale BNNBase BNN at s₁Generating function

Same generated sample; feature dimension d = 1. Curves illustrate the construction for this sample, rather than a performance or coverage comparison.

a1 + a2 = 1,   a1s1−2/d + a2s2−2/d = 0
μ̂two-scale = a1μ̂s₁ + a2μ̂s₂.

For s₁ < s₂, the first coefficient is negative. This is how the shared leading term cancels. Predictions can consequently leave the range of observed responses. Bringing the scales too close makes the coefficients large; finite-sample stability remains part of scale selection.

Theory and conditions · Exact coefficients and algorithm.