Models and their options¶
SPDnet, RResNet and GBWBNRResNet take 28 to 37 constructor arguments,
because they expose every option of the layers they contain. Most can be
left at their default. This page groups them by role; the
Models page lists them one by one.
Which model¶
SPDnetThe SPDNet of Huang & Van Gool (AAAI 2017):
[BiMap → ReEig → (BatchNorm)] × L → LogEig → Vec → Linear, withL = len(hidden_layers_size). The default choice.RResNetA Riemannian residual network (Katsman et al., NeurIPS 2023) with several stages:
[BiMap → (ReEig) → (BatchNorm) → ResidualBlock × k] × L. The residual blocks move each matrix along a learned tangent vector (Riemannian residual networks). No ReEig by default.GBWBNRResNetThe single-stage residual network of the GBWBN paper (Wang et al., 2025):
BiMap → GBWBN → ResidualBlock → LogEig. Kept to reproduce that paper;RResNetcovers the same architecture with more options.
Arguments by role¶
Shared by the three models unless noted.
Role |
Arguments |
What to know |
|---|---|---|
Sizes |
|
|
Non-linearity |
|
Eigenvalues below |
BiMap weights |
|
Stiefel constraint, static or dynamic (Layers and parametrizations). |
Batch normalization: which |
|
|
Batch normalization: statistics |
|
Running statistics and smoothing of the batch mean. |
Batch normalization: bias |
|
How the SPD bias stays SPD. |
GBWBN |
|
|
Residual blocks (RResNet, GBWBNRResNet) |
|
The spectral vector field and the exponential map (Riemannian residual networks). |
Head |
|
|
Numerics |
|
|
After building the optimizer¶
If any part uses a dynamic parametrization (*_parametrization_mode="dynamic"),
register the optimizer once:
optimizer = torch.optim.SGD(model.parameters(), lr=0.05, momentum=0.9, nesterov=True)
model.register_optimizer_hook(optimizer)
Configurations of the published experiments¶
These configurations reproduce published results with this library (see the
spdnet-benchmark-demo repository for the training loop).
SPDNet batch normalization paper, HyperLeaf (204×204 hyperspectral covariances, 4 cultivars); SGD with Nesterov momentum, 5 warm-up epochs, learning rate halved on plateau, early stopping:
SPDnet(
input_dim=204,
hidden_layers_size=[184, 158],
output_dim=4,
reeig_eps=0.01,
batchnorm=True,
batchnorm_mean_type="arithmetic", # batch 48, lr 0.05
)
Same paper, HDM05 (93×93 skeleton covariances scaled by 190, 117 classes):
SPDnet(
input_dim=93,
hidden_layers_size=[84, 63],
output_dim=117,
reeig_eps=0.8,
batchnorm=True,
batchnorm_mean_type="geometric_arithmetic_harmonic", # batch 16, lr 0.25
)
GBWBN paper, HDM05 (raw matrices; Adam with AMSGrad, lr 2.5e-3, batch 30, 200 epochs):
GBWBNRResNet( # GBWBN (momentum 0.1) is on by default
input_dim=93,
hidden_dim=30,
output_dim=117,
batchnorm_bw_options={"bw_theta": 0.5},
)