Skip to content

Benchmarks ​

Lazy normalization preserves the original matrix representation. Eager normalization pays a construction cost and then operates on normalized dense storage. Which choice helps depends on matrix size, sparsity, backend, and the operations your algorithm repeats.

Initial fitting measurements ​

These measurements compare a Gaussian ridge fit through lazy and eager operators, with an intercept and penalty of 0.1. Both paths fit the same problem, and the harness checks coefficients, intercepts, and objectives after converting eager coefficients back with the fitted parameters.

Exploratory measurements

These are single instrumented runs, collected while other builds were active. They illustrate the tradeoffs and are not stable performance rankings. Run the benchmarks on your target hardware before choosing a representation.

Lazy fit3.16 ms
Eager build + fit7.55 ms
Gaussian ridge fit with standardized predictors. Eager time includes dense construction. Single instrumented runs from October 8, 2026.
Measured phaseTime (ms)Rust allocationsPeak extra memory (KiB)
Lazy fit3.166772.0
Eager construction1.8082068.0
Eager fit5.75969.0

The eager output alone occupies 2048.0 KiB and remains live during fitting. The eager-fit memory figure excludes this matrix because it was allocated before that phase.

Download all raw measurements, including the centered-only cases. The chart and table are generated from the checked-in CSV.

Read memory figures across phases ​

The allocator records the peak live requested Rust bytes above the live bytes at the start of each phase. Eager construction includes its dense output, but eager fitting starts with that output already allocated. A lower eager-fit peak therefore does not imply lower overall memory.

The measurements exclude native BLAS/LAPACK allocations, allocator overhead, and resident-memory effects. Total process memory also depends on whether the original input stays live. For owned writable dense inputs, into_eager can normalize the original storage instead of creating a second matrix.

Separate setup from repeated work ​

For k repetitions of a fixed operation, compare

Tlazy(k)=Tfit+ktlazy,Teager(k)=Tfit+Tmaterialize+kteager.

When teager<tlazy, materialization pays for itself after roughly

k>Tmaterializetlazy−teager.

This estimate assumes the same fitted parameters and fixed operation costs. The fitting measurements above cover complete solver runs, whose operation counts can differ slightly. They should not be used to infer a per-product break-even point. Sparse lazy products can also remain faster than dense eager products.

Methodology and reproduction ​

The allocation profiles were collected on October 8, 2026, with Rust 1.89.0 in release mode on an Intel Core Ultra 7 155U. Fixtures use seed 205, sizes 512×32 and 2048×128, and target densities of 100% and 10%. OpenBLAS and OpenMP use one thread. Fixture generation and source-storage conversion occur outside measurement, and each fitting path is warmed first.

sh
# Construction and repeated forward/transpose products with Criterion.
cargo bench --locked --bench eager_normalization --features ndarray,sprs

# Consumer fitting times and Rust allocation profiles.
bash scripts/bench-normalization-allocations.sh > target/normalization-allocations.csv

The measurement notes describe the initial runs. The Criterion harness separates parameter fitting, dense materialization, and batches of forward and transpose products. Use its longer default measurement settings on an idle machine for runtime comparisons.

Released under the MIT and Apache 2.0 licenses.