Layers¶
yetanotherspdnet.nn — the SPDNet layers. BiMap changes the dimension,
ReEig / ReEigBias are the non-linearities, LogEig and Vec / Vech map
to a vector for a Euclidean head. SampleCovariance and MEstimation build
SPD matrices from raw samples, differentiably. See Layers and parametrizations
for the equations and the static / dynamic parametrization of BiMap.
Bilinear mapping layer \(X \mapsto W^\top X W\) reducing SPD dimension. |
|
Eigenvalue rectification \(X \mapsto U \max(\Lambda, \epsilon I) U^\top\). |
|
Eigenvalue rectification with a learnable shift of each eigenvalue. |
|
Matrix logarithm \(X \mapsto U \log(\Lambda) U^\top\). |
|
Full vectorization of a batch of matrices: |
|
Half vectorization of symmetric matrices: |
|
Sample covariance matrix of a batch of sample sets. |
|
Robust covariance estimation layer (M-estimator of scatter). |
SPDNet layers¶
- class BiMap(
- n_in,
- n_out,
- parametrized=True,
- parametrization_mode='static',
- parametrization_options=None,
- n_steps_ref_update=100,
- use_autograd=False,
- device=device(type='cpu'),
- dtype=torch.float64,
- generator=None,
Bilinear mapping layer \(X \mapsto W^\top X W\) reducing SPD dimension.
With \(W \in \mathbb{R}^{n_{in} \times n_{out}}\) on the Stiefel manifold (\(W^\top W = I\)), a congruence maps SPD matrices to SPD matrices of smaller size. Introduced in Huang & Van Gool, A Riemannian Network for SPD Matrix Learning, AAAI 2017.
Build a BiMap layer.
- Parameters:
n_in (
int) – Number of input featuresn_out (
int) – Number of output featuresparametrized (
bool, optional) – Whether to apply parametrization to enforce orthogonal constraint. Default is Trueparametrization_mode (
str, optional) – Parametrization mode. Default is “static”. Choices are: “static” and “dynamic”parametrization_options (
dict, optional) – Options for the parametrization function. Default is Nonen_steps_ref_update (
int, optional) – If parametrization_mode is “dynamic”, number of steps in between each reference point update. Default is 100use_autograd (
bool, optional) – Use torch autograd for the computation of the gradient rather than the analytical formula. Default is Falsedevice (
torch.device, optional) – Device on which the layer is initialized. Default is torch.device(“cpu”)dtype (
torch.dtype, optional) – Data type of the layer. Default is torch.float64generator (
torch.Generator, optional) – Generator to ensure reproducibility. Default is None
- Variables:
weight (
torch.Tensorofshape (n_in,n_out)) – Orthonormal weight. Whenparametrizedit is atorch.nn.utils.parametrizeview whose trainable tensor isparametrizations.weight.original.n_out (n_in,) – Input and output matrix dimensions (
n_out <= n_in).is_dynamic (
bool) – Whether the Stiefel parametrization re-centres its reference point everyn_steps_ref_updateoptimizer steps (parametrization_mode="dynamic").use_autograd (
bool) – Autograd path (True) or hand-written backward (False, default).
- forward(data)[source]¶
Forward pass of the BiMap layer
- Parameters:
data (
torch.Tensorofshape (...,n_in,n_in)) – Batch of SPD matrices- Returns:
data_transformed (
torch.Tensorofshape (...,n_out,n_out)) – Batch of transformed SPD matrices- Return type:
- register_optimizer_hook(optimizer)[source]¶
Register the post-step hook with the optimizer. If dynamic parametrization, it needs to be called once after creating the optimizer for dynamic parametrization to actually work as expected
- Parameters:
optimizer (
torch.optim.Optimizer) – Torch optimizer used for training
- class ReEig(eps=0.0001, use_autograd=False, dim=None)[source]¶
Eigenvalue rectification \(X \mapsto U \max(\Lambda, \epsilon I) U^\top\).
The SPD analogue of ReLU: eigenvalues below
epsare clamped, which keeps the output well conditioned. Non-linear on the manifold, it plays the role of the activation in SPDNet.- ReEig layer in a SPDnet layer according to the paper:
A Riemannian Network for SPD Matrix Learning, Huang et al AAAI Conference on Artificial Intelligence, 2017
- Parameters:
eps (
float, optional) – Value of rectification of the eigenvalues. Default is 1e-2.use_autograd (
bool, optional) – Use torch autograd for the computation of the gradient rather than the analytical formula. Default is False.dim (
int, optional) – Dimension of the SPD matrices. Default is None. Used for logging purposes.
- Variables:
- forward(data)[source]¶
Forward pass of the ReEig layer
- Parameters:
data (
torch.Tensorofshape (...,n_features,n_features)) – Batch of SPD matrices- Returns:
data_transformed (
torch.Tensorofshape (...,n_features,n_features)) – Batch of transformed SPD matrices- Return type:
- class ReEigBias(
- n_features,
- eps=0.0001,
- use_autograd=False,
- device=device(type='cpu'),
- dtype=torch.float64,
Eigenvalue rectification with a learnable shift of each eigenvalue.
\[X \mapsto U \operatorname{diag}\big( \operatorname{clamp}(\lambda_i + b_i,\ \epsilon,\ 1/\epsilon)\big) U^\top\]with eigenvalues in ascending order and \(b\) learned (initialized at zero, so the layer starts as a two-sided ReEig). The upper bound \(1/\epsilon\) caps the condition number of the output, which is useful after estimators such as M-estimators that can produce large eigenvalues.
Build a ReEigBias layer.
- Parameters:
n_features (
int) – Dimension of the SPD matrices (size of the bias vector)eps (
float, optional) – Lower clamping value of the shifted eigenvalues; the upper one is1 / eps. Default is 1e-4use_autograd (
bool, optional) – Use torch autograd for the computation of the gradient rather than the analytical formula. Default is Falsedevice (
torch.device, optional) – Device on which the layer is initialized. Default is torch.device(“cpu”)dtype (
torch.dtype, optional) – Data type of the layer. Default is torch.float64
- Variables:
- forward(data)[source]¶
Forward pass of the ReEigBias layer
- Parameters:
data (
torch.Tensorofshape (...,n_features,n_features)) – Batch of symmetric matrices- Returns:
data_transformed (
torch.Tensorofshape (...,n_features,n_features)) – Batch of SPD matrices- Return type:
- class LogEig(use_autograd=False)[source]¶
Matrix logarithm \(X \mapsto U \log(\Lambda) U^\top\).
Maps SPD matrices to the (flat) space of symmetric matrices, usually right before vectorization and a Euclidean classifier.
- LogEig layer in a SPDnet layer according to the paper:
A Riemannian Network for SPD Matrix Learning, Huang et al AAAI Conference on Artificial Intelligence, 2017
- Parameters:
use_autograd (
bool, optional) – Use torch autograd for the computation of the gradient rather than the analytical formula. Default is False.- Variables:
use_autograd (
bool) – Autograd path (True) or hand-written Daleckii-Krein backward (False).
- forward(data)[source]¶
Forward pass of the LogEig layer
- Parameters:
data (
torch.Tensorofshape (...,n_features,n_features)) – Batch of SPD matrices- Returns:
data_transformed (
torch.Tensorofshape (...,n_features,n_features)) – Batch of transformed symmetric matrices- Return type:
- class Vec(use_autograd=False)[source]¶
Full vectorization of a batch of matrices:
(..., n, n) -> (..., n * n).Vectorization operator of a batch of matrices according to the last two dimensions
- Parameters:
use_autograd (
bool, optional) – Use torch autograd for the computation of the gradient rather than the analytical formula. Default is False.- Variables:
use_autograd (
bool) – Autograd path (True) or hand-written backward (False).
- forward(data)[source]¶
Forward pass of the Vec layer
- Parameters:
data (
torch.Tensorofshape (...,n_rows,n_columns)) – Batch of matrices- Returns:
data_vec (
torch.Tensorofshape (...,n_rows*n_columns)) – Batch of vectorized matrices- Return type:
- class Vech[source]¶
Half vectorization of symmetric matrices:
(..., n, n) -> (..., n (n + 1) / 2).Keeps only the upper-triangular part, which carries all the information of a symmetric matrix.
Vech operator of a batch of matrices according to the last two dimensions
WARNING : no automatic differentiation available here because it fails
- forward(data)[source]¶
Forward pass of the Vech layer
WARNING : no automatic differentiation available here because it fails
- Parameters:
data (
torch.Tensorofshape (...,n_features,n_features)) – Batch of symmetric matrices- Returns:
data_vech (
torch.Tensorofshape (...,n_features*(n_features+1)//2)) – Batch of vech matrices- Return type:
Covariance estimation¶
- class SampleCovariance(assume_centered=False)[source]¶
Sample covariance matrix of a batch of sample sets.
Maps samples
(..., n_samples, n_features)to SPD matrices(..., n_features, n_features), typically as the first layer of an SPD network fed with raw signals.Build a SampleCovariance layer.
- Parameters:
assume_centered (
bool, optional) – Whether the samples are already centered. Default is False- Variables:
assume_centered (
bool) – Whether the samples are assumed centered.
- forward(data)[source]¶
Forward pass of the SampleCovariance layer
- Parameters:
data (
torch.Tensorofshape (...,n_samples,n_features)) – Samples- Returns:
covariance (
torch.Tensorofshape (...,n_features,n_features)) – Sample covariance matrices- Return type:
- class MEstimation(
- weight_function,
- n_iterations=30,
- tol=1e-06,
- assume_centered=False,
- shrinkage=None,
- normalize=None,
- use_autograd=False,
Robust covariance estimation layer (M-estimator of scatter).
See
yetanotherspdnet.functions.m_estimatorsfor the estimator and its two gradient paths. The layer has no trainable parameter; it makes the estimator usable in atorch.nn.Sequentialand differentiable with respect to its input.Build an MEstimation layer.
- Parameters:
weight_function (
Callable) – Weight of squared Mahalanobis distances, e.g.functools.partial(student_function, n_features=p, nu=3.0)n_iterations (
int, optional) – Maximum number of fixed-point iterations. Default is 30tol (
float, optional) – Stopping tolerance on the relative change. Default is 1e-6assume_centered (
bool, optional) – Whether the samples are already centered. Default is Falseshrinkage (
float | None, optional) – Shrinkage coefficient towards the identity. Default is Nonenormalize (
str | None, optional) –"trace","determinant"or None, applied at each iteration; required for scale-invariant weights such as Tyler’s. Default is Noneuse_autograd (
bool, optional) – Unrolled autograd path (True) or implicit backward (False, default)
- Variables:
weight_function (
Callable) – Weight function of the estimator.tol (n_iterations,) – Stopping criteria of the fixed-point iterations.
use_autograd (
bool) – Gradient path.
- forward(data, init=None)[source]¶
Forward pass of the MEstimation layer
- Parameters:
data (
torch.Tensorofshape (...,n_samples,n_features)) – Samplesinit (
torch.Tensor, optional) – Initial estimate. Default is the identity
- Returns:
covariance (
torch.Tensorofshape (...,n_features,n_features)) – Estimated scatter matrices- Return type: