How the library fits together¶
This page is the map: what the three levels of the package do, what flows between them, and the conventions every function and layer follows. The other guides go into each part.
Three levels¶
model SPDnet, RResNet, GBWBNRResNet classifiers, one constructor
│ builds a torch.nn.Sequential of ↓
nn BiMap, ReEig, LogEig, BatchNorm…, ResidualBlock modules with parameters
│ call ↓ and keep their parameters on a manifold with parametrizations
functions spd_linalg, spd_geometries/*, m_estimators, stiefel pure tensor functions
Dependencies only go downwards. You can use any level on its own: compute a
Karcher mean with functions.spd_geometries.affine_invariant, put a
BatchNormSPDMean in your own network, or train an SPDnet directly.
What flows through a network¶
A batch of SPD matrices has shape (B, n, n); a batch of sequences of SPD
matrices (B, L, n, n) (every function works on the last two dimensions).
Following one SPDnet on HDM05 (93×93 skeleton covariances, 117 classes):
Step |
Operation |
Shape |
|---|---|---|
input |
covariance of the joint coordinates |
|
|
\(X \mapsto W^\top X W\), \(W \in \mathrm{St}(93, 84)\) |
|
|
\(U \max(\Lambda, \epsilon) U^\top\) |
|
|
centre the batch at a learned SPD bias |
|
|
same, \(84 \to 63\) |
|
|
\(U \log(\Lambda) U^\top\): symmetric, flat space |
|
|
flatten (\(n^2\) or \(n(n+1)/2\) entries) |
|
|
Euclidean classifier |
|
Everything before LogEig stays on the SPD manifold; the layers are designed
so that their output is SPD whenever their input is.
Conventions¶
float64 by default. Eigendecompositions of ill-conditioned matrices lose
accuracy quickly in float32, and the gradients of spectral functions divide
by eigenvalue gaps. Every layer and function takes dtype and device
arguments; float32 is possible but not the tested default.
Two gradient paths. Every operation built on an eigendecomposition
exists as a snake_case function differentiated by autograd and a
CamelCase torch.autograd.Function with a hand-written backward. Layers
choose with use_autograd, which defaults to False (hand-written). The
hand-written backwards stay finite and exact when eigenvalues are close or
equal, which autograd through torch.linalg.eigh does not
(Numerical and optimization techniques).
Matrix functions return tuples. logm_SPD, sqrtm_SPD, powm_SPD, …
return (result, eigvals, eigvecs) so that callers can reuse the
decomposition; take [0] for the matrix. The Function classes return the
matrix only.
Parameters live on manifolds. BiMap weights must stay orthonormal and
batch normalization biases SPD. Both are torch.nn.utils.parametrize
parametrizations: the optimizer updates an unconstrained tensor, and a map
sends it to the manifold. The map can be fixed (static) or re-centred
during training (dynamic) (Layers and parametrizations).
Where to go next¶
To understand… |
Read |
|---|---|
BiMap, ReEig, LogEig and the static / dynamic parametrizations |
|
which geometry to use and what its mean is |
|
what the batch normalization does, and its options |
|
the residual block of RResNet |
|
the constructor arguments of the models, grouped by role |
|
|