BNNBagged nearest
neighbors
Nonparametric estimation by averaging nearest neighbors over subsamples.
Start with a simple prediction rule. Average it over possible subsamples. The resulting estimator has an exact, directly computable representation.
A distribution
over neighbors
For a query x, draw s observations without replacement and select the closest one. BNN averages that observation's response across every possible size-s subsample.
Once the full sample is ordered by distance, each response receives a known weight. This connects a statistical construction to a practical algorithm: distances, an ordering, and a weighted sum.
Locality has a clear control
The subsample size determines the neighborhood. With one observation per subsample, the estimate is the sample mean. With the whole sample in each subsample, it is the ordinary nearest-neighbor response. Intermediate scales spread influence across the distance ranks.
A larger subsample is more likely to contain a close observation. Increasing the scale therefore makes the estimate more local. The interactive explanation lets you inspect this relationship in both the weights and the fitted curve.
- Exact computationThe average over all subsamples is calculated through rank weights, without enumerating the subsamples.
- Bias correctionA two-scale combination can cancel a leading bias term. Its signed weights and statistical conditions are explained explicitly.
- Statistical inferenceDistribution theory, resampling, uniform results, and transformed estimands answer different questions about estimation error.
From the construction to its properties
The estimator's rank representation provides a route into its statistical behavior. The theory includes bias expansions, pointwise asymptotic distributions, and resampling inference. Further results concern control over a region, derivatives, quotients, and bootstrap transformations.
The theory guide connects each question to its result and source. It keeps the evaluation domain, centering, scale growth, and additional conditions visible alongside the conclusion.
Read the statistical theory ↗
Exact identities, asymptotic results, and what uniformity means.
Follow the algorithms ↗
Prediction, inference, reuse of calculations, and measured workloads.
Inspect a regression example ↗
A reproducible illustration with its generating function and code.
Run the software ↗
NumPy, PyTorch, and JAX implementations in one research repository.
Terminology and sources
This site uses BNN for the fixed-size, without-replacement bagging construction. The statistical papers use distributional nearest neighbors (DNN), and two-scale DNN (TDNN) for its bias-corrected extension. These abbreviations concern nearest-neighbor estimation.
The papers and references credit the earlier foundations and the work developing estimation and inference. Algorithm definitions, numerical checks, examples, and source provenance accompany the maintained implementation.