Theory

The statistical questions,
made explicit.

An exact estimator, its bias and distribution, and the additional questions that arise when inference ranges over a region or targets a transformed function.

An exact finite-sample object

Let (Xᵢ,Yᵢ), i = 1,…,n, be the observations. At a query x, order their responses by increasing distance and write Y(i) for the response at rank i. BNN averages the nearest response over all size-s subsamples drawn without replacement.

μ̂s(x) = C(n,s)−1 Σ|I|=s Ynearest in I
= Σi=1n [C(n−i,s−1)/C(n,s)] Y(i).

This equality is combinatorial. It holds for the realized sample and specified tie rule. Distributional assumptions enter when studying how the estimator behaves as the sample changes.

The weights are nonnegative, sum to one, and depend only on n and s after ordering. The U-statistic representation supports probabilistic analysis; the rank representation supports computation.

Steele (2009); JASA Section 2, Lemma 1 and equations (5)–(6).

Bias depends on scale and dimension

Consider Y = μ(X) + ε and a fixed query in the support. Under the inference paper's smoothness, design, moment, and sampling conditions, the base estimator's expectation has a scale-dependent expansion.

E[μ̂s(x)] − μ(x) = b(x)s−2/d + R(s).
R(s) = O(s−4/d) for d ≥ 2; O(s−3) for d = 1.

The coefficient b(x) depends on the local regression function and feature density, but not on s. Its sign depends on that local structure. As s grows, the estimator becomes more local. Sampling variability and locality must be balanced rather than optimized separately.

Conditions and interpretation

JASA Conditions 1–3 specify sampling, tail behavior, and the local density and regression smoothness. The expansion uses four derivatives, a density bounded away from zero and infinity near the query, and moment and noise conditions. Feature dimension d is fixed.

A scale can grow while remaining small relative to n. The conditions control both the approach to the target and the randomness of the estimate. The exact finite-sample identity above does not require these conditions.

JASA Theorem 1. Inspect locality in the explorer.

Bias cancellation and convergence

The shared coefficient b(x) allows two estimates to be combined. Choose a₁+a₂=1 and a₁s₁−2/d+a₂s₂−2/d=0. Constants are preserved and the leading bias term cancels.

μ̂TDNN(x) = a1μ̂s₁(x) + a2μ̂s₂(x),
a1 = −1 / [(s2/s1)2/d−1],   a2 = 1−a1.

The two scales grow and their ratio stays bounded away from zero and one. Under the paper's conditions, the remainder and variance yield the pointwise error-rate analysis. For d ≥ 2, balancing the squared-bias term s−8/d against a variance term of order s/n gives the rate n−4/(d+8) for the estimator's error scale. The paper treats d = 1 separately.

The combination changes the weights

Negative weights permit cancellation. They also affect variability: two individually asymptotically normal estimates need a joint argument to justify their combined distribution. The source develops that argument for the two-scale estimator.

JASA Theorems 3–4, with the scale and smoothness conditions in the article. Computational coefficients.

Distribution and uncertainty at a query

Pointwise asymptotic normality describes the estimator at a fixed x after appropriate centering and normalization. Bias centering and studentization are part of the statement.

[μ̂s(x) − μ(x) − B(s,x)] / σn(x) → N(0,1).

The base and two-scale results have their own bias and variance terms. A confidence interval centered directly on the estimate needs the residual bias to be negligible on the relevant standard-error scale, or an appropriate correction.

JASA develops jackknife and bootstrap approaches to variance estimation, and its supplement addresses bootstrap distribution approximation and treatment-effect inference. Conditions for these inference procedures can be stronger than those for a prediction or bias expansion.

What the implementation holds fixed

The NumPy bootstrap resamples n paired observations and recomputes the fixed-scale estimator. Draws are paired across queries. Its sample variance uses denominator B−1. The delete-one jackknife follows the paper's original-estimate-centered formula and requires the scales to be valid after deletion.

These procedures do not automatically include uncertainty introduced by a separate scale-selection rule. The algorithm guide specifies those choices.

JASA Theorems 2–3, Section 4 and supplement.

Uniform consistency over a region

A pointwise statement concerns one query. Uniform consistency concerns the largest error over an entire evaluation region K. This matters when a procedure uses many locations or nearby queries.

supx ∈ K |μ̂s(x) − μ(x)| →p 0.

Wang and Huang's final appendix states a uniform-convergence result in Theorem C.3. Its setup includes independent sampling, compact support, regular density and regression smoothness, bounded-kernel and noise conditions, and growing subsamples with s/n → 0. The appendix uses m for the scale denoted s here.

Uniform analysis must control how estimation errors fluctuate between locations. Compactness and pointwise convergence alone do not supply that control. A grid argument needs control between grid points; an empirical-process argument needs control of the estimator class.

Read the source statement and assumptions

Theorem C.3 uses Assumptions C.1–C.5. Section C.6 presents the concentration and stochastic-equicontinuity arguments. The statement concerns the support specified in that setup; derivative calculations also require nearby evaluation points to be admissible.

Theorem C.3 in the final online appendix. This guide summarizes the source's statistical objects and conditions; the cited appendix contains its full statements and proofs.

Wang and Huang, final online appendix, Section C.3, Theorem C.3, dated July 7, 2026.

Uniform normal approximation

Theorem C.4 states a normal approximation controlled across query locations and distribution thresholds. Its object is the largest difference between the standardized estimator's marginal CDF and the normal CDF.

supx ∈ K supt ∈ ℝ |P(Tn(x) ≤ t) − Φ(t)| → 0.

Here Tn includes the source's centering and variance normalization. The source adds variance nondegeneracy to its regularity and scale conditions. Uniformity makes the approximation comparable over a region rather than just at one chosen query.

Uniform marginal inference and simultaneous bands

A bound on marginal distribution approximation is different from the distribution of supx∈K|Tn(x)|. A simultaneous band requires the latter object, joint dependence across queries, and the relevant approximation and anti-concentration arguments.

The final appendix explicitly distinguishes its uniform pointwise confidence-interval treatment from a formally established simultaneous confidence-band result.

Theorem C.4; discussion following the proof of Theorem C.8 in the final appendix.

Derivatives and quotients need additional control

The appendix extends the statistical discussion to finite-difference derivative estimation (C.5), joint scale/step-size analysis (C.6), and quotients (C.7). These operations explain why controlling the regression surface over a region is useful.

Derivative estimation

A centered difference evaluates nearby predictions:

Dhμ̂(x) = [μ̂(x+h ej/2) − μ̂(x−h ej/2)] / h.

If the regression error is bounded by rn on a region containing both shifted queries, and μ has a bounded third derivative there, an elementary error bound is

|Dhμ̂(x) − ∂jμ(x)| ≤ 2rn/h + C h².

This shows the two roles of the step: a smaller h reduces the approximation error while amplifying regression error. A direct consistency argument needs a compatible relation between rn and h, or sharper control of increments. The appendix's stated derivative and rate results should be read with their specific difference convention and assumptions.

Theorems C.5–C.6. The displayed centered-difference inequality explains the operation; it does not restate an optimal-rate claim.

Quotient estimation

For a ratio μ₁(x)/μ₂(x), uniform consistency of both inputs can transfer to the ratio when the denominator stays separated from zero. The denominator condition is substantive: a small absolute regression error can produce a large ratio error near zero.

μ̂₁ / μ̂₂ →p μ₁ / μ₂ uniformly on K,
with uniform input control and infx∈K|μ₂(x)| > 0.

Theorem C.7 and the appendix's continuous-mapping lemma.

Bootstrap for a family of estimands

The appendix's bootstrap results address the normalized estimator and then transformed objects. They give the theory page a broader scope than point prediction alone.

Results in the final appendix
ResultStatistical objectWhat to inspect
C.8Uniform bootstrap normal approximationConditional centering, studentization, growing scales, and nondegeneracy
C.9Uniform approximation to the sampling distributionHow the bootstrap and original standardized CDFs are compared
C.10Finite-difference estimatorsShrinking step size, increments, centering, and the transformed distribution
C.11Transformations of levels and derivativesJoint process behavior, differentiability, and variance normalization

Passing from marginal convergence to a function-valued process or a changing finite-difference map requires the corresponding process and rate arguments. The stated assumptions and transformations belong with the result.

Wang and Huang, final online appendix, Sections C.5–C.6. Paper and full attribution.

Read a result with its conditions

The exact weighted-sum identity, a bias expansion, a normal approximation, and a bootstrap theorem are different claims. Check the evaluation domain, smoothness, moments, dimension, scale growth, residual bias, variance normalization, and any transformation-specific restrictions.

Numerical tests check finite-sample calculations. A regression illustration displays a realized fit. A simulation estimates performance under a specified generating model. The papers supply the asymptotic statements and arguments. Read the formal sources or follow the computational routines.