BCESiameseNetwork¶
Both images are passed through the same shared-weight backbone and
projection head. Each branch output is squashed with a sigmoid into a
feature vector h in (0, 1)^D. The two branches are then combined by their
component-wise L1 distance, and a single learned linear layer maps that
distance vector to the probability of the pair showing the same class:
where the weights \(\alpha_j\) learn the importance of each feature
dimension, so unlike ContrastiveSiameseNetwork the comparison
metric itself is trained. The network is a binary classifier over pairs and
is trained with binary cross-entropy on labels 1 (same class) / 0
(different class).
Following diagram visualizes this:
┌──────────┐ ┌────────────────┐ ┌─────────┐
Input Image A ───►│ Backbone │───►│ Embedding Head │───►│ Sigmoid │───► Features A ──┐
└──────────┘ └────────────────┘ └─────────┘ │
╎ ╎ ╎ │ ┌─────────┐ ┌───────────────┐
╎ ╎ ╎ ├───►│ |A - B| │───►│ Scoring Layer │───► P(same class)
╎ ╎ ╎ │ └─────────┘ └───────────────┘
┌──────────┐ ┌────────────────┐ ┌─────────┐ │
Input Image B ───►│ Backbone │───►│ Embedding Head │───►│ Sigmoid │───► Features B ──┘
└──────────┘ └────────────────┘ └─────────┘
(Shared Weights)
Example: training a BCE Siamese Network¶
See this tutorial.
API reference¶
- class pyvisim.neural_networks.BCESiameseNetwork(backbone='resnet18', embedding_dim=128, transform=None, device='cpu', pretrained_backbone=True, *, batch_size=16)[source]¶
Bases:
BackboneWithHeadSiamese network that classifies image pairs, proposed in Koch, G., Zemel, R., & Salakhutdinov, R. (2015). Siamese Neural Networks for One-shot Image Recognition.
For more information, see the documentation:
https://mechacritter.github.io/Python-Visual-Similarity/neural_networks/bce_siamese/bce_siamese.html.NOTE¶
The score is a learned probability, not a geometric similarity: it is symmetric in its inputs (the L1 distance is), lives in
(0, 1), and for two identical images equalssigmoid(b)– the learned bias sets the operating point, so a perfect match does not score exactly1.References:¶
[1] Koch, G., Zemel, R., & Salakhutdinov, R. (2015). Siamese Neural Networks for One-shot Image Recognition. ICML Deep Learning Workshop. https://www.cs.cmu.edu/~rsalakhu/papers/oneshot1.pdf
- param backbone:
name of feature-extraction network. See
https://mechacritter.github.io/Python-Visual-Similarity/neural_networks/backbones/backbones.html.- param embedding_dim:
Dimensionality of the twin feature vectors that the scoring layer compares.
- param transform:
processing transform applied to every input image. If
None, the ImageNet preprocessing matching the backbone is used.- param device:
Device on which the model is placed.
- param pretrained_backbone:
Whether to use a backbone pretrained on ImageNet. If you are loading the
BCESiameseNetworkfrom a checkpoint, set this toFalseto avoid downloading the weights again.- param batch_size:
Maximum number of images processed in a single batch. Set to
-1to process all images as a single batch.- raises ValueError:
If
embedding_dimis not a positive integer or ifbackboneis not a supported backbone name.
- embed(images, *, dims='HWC', value_range=(0.0, 255.0))[source]¶
Not implemented for this class. Please do not use!
- Parameters:
images (_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes] | Iterable[_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes]])
dims (str)
- Return type:
- forward(x1, x2)[source]¶
Computes same-class logits for a batch of aligned image pairs.
The i-th logit scores the pair
(x1[i], x2[i]); applytorch.sigmoid()to obtain probabilities, or feed the logits directly totorch.nn.BCEWithLogitsLossduring training.- Parameters:
- Returns:
Logit tensor of shape (batch,).
- Raises:
ValueError – If the two batches differ in shape.
- Return type:
- classmethod from_dict(state, **kwargs)¶
Rebuilds the embedder a state dictionary describes.
Called on
SerializableImageEmbedderitself, it hands the state to thefrom_dictof the class named under"__class__", so a state can be rebuilt without knowing which embedder wrote it.- Parameters:
- Returns:
The reconstructed embedder.
- Raises:
ValueError – If
statenames no concrete embedder class.NotImplementedError – If called on a subclass that does not implement its own
from_dict.
- Return type:
_NeuralEmbedderT
- classmethod load_from_disk(path, **kwargs)¶
Loads an object previously saved with
save_to_disk().Not every part of an object survives serialization: an arbitrary callable such as a torchvision transform has no portable description, so it is left out of the file. Pass such an object back here as a keyword argument.
- Parameters:
kwargs (Any) – Objects the file cannot hold, forwarded to
from_dict().
- Returns:
A ready-to-use instance.
- Raises:
FileNotFoundError – If
pathdoes not exist.ValueError – If the file is not a valid file of this kind or was saved by a different class.
TypeError – If the class does not take one of
kwargs.
- Return type:
_SerializableT
- save_to_disk(path)¶
Saves the serialized state of this object to a file.
- set_batch_size(batch_size)¶
Sets the number of items processed per batch.
- Parameters:
batch_size (int) – Maximum number of images processed in a single batch. Set to
-1to process all images as a single batch.- Raises:
ValueError – If
batch_sizeis neither-1nor a positive integer.- Return type:
None
- similarity_score(images1, images2, *, dims='HWC', value_range=(0.0, 255.0))[source]¶
Compute the similarity scores matrix between two (batches of) images.
- Parameters:
images1 (_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes] | Iterable[_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes]]) – First (batch of) image(s) as
MatLike(NumPy array, torch tensor or array-like).images2 (_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes] | Iterable[_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes]]) – Second (batch of) image(s) as
MatLike.dims (str) – Axis-label string, one character per array axis in order:
"H"= height (rows),"W"= width (columns),"C"= channels (e.g. RGB),"B"= batch size. For example,"HWC"is height × width × channels (NumPy/OpenCV single-image layout);"CHW"is channels × height × width (PyTorch single-image layout);"BCHW"is batch × channels × height × width (PyTorch batched layout). Seepyvisim.typing.value_range (tuple[float, float]) – The
(low, high)range the input values live in; converted into the canonical[0, 255]range.
- Returns:
The similarity score matrix of shape
(len(images1), len(images2)).- Return type:
- to_dict()¶
Serializes this object into a JSON-safe state dictionary.
The mapping holds the output of
_state()plus the format version under"format_version"and the class name under"__class__". Arrays may be embedded as__ndarray__nodes, which the serialization layer stores as binary tensors.- Returns:
A JSON-safe description suitable for
from_dict().- Return type:
- property device: device¶
The device the model’s parameters live on.
Derived from the parameters themselves rather than cached, so it stays correct after the user moves the model with
model.to(...).