SIFT¶
Scale-Invariant Feature Transform descriptors. SIFT was the original local descriptor used for VLAD and Fisher Vector embedding.
output_dimis128(standard SIFT descriptor length).
For most uses prefer RootSIFT, which normalizes these descriptors and usually improves retrieval at no extra cost.
References¶
D. G. Lowe. “Distinctive Image Features from Scale-Invariant Keypoints”. In: International Journal of Computer Vision 60.2 (2004), pp. 91-110. issn: 1573-1405. doi: 10.1023/B:VISI.0000029664.99615.94. url: https://doi.org/10.1023/B:VISI.0000029664.99615.94.
API reference¶
- class pyvisim.features.SIFT(upsampling=2, n_octaves=8, n_scales=3, sigma_min=1.6, sigma_in=0.5, c_dog=0.013333333333333334, c_edge=10, n_bins=36, lambda_ori=1.5, c_max=0.8, lambda_descr=6, n_hist=4, n_ori=8)[source]¶
Bases:
FeatureExtractorBase,SIFTScale-Invariant Feature Transform (SIFT) feature extractor.
References:¶
[1] Lowe, D. G. (2004). Distinctive image features from scale-invariant keypoints.
- __call__(image, /, *, dims='HWC', value_range=(0.0, 255.0))[source]¶
Extracts features from an image.
- Parameters:
image (_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes]) – Input image as
MatLike(NumPy array, torch tensor or array-like). It is normalized to a canonicaluint8(H, W, C)image before extraction.dims (str) – Axis-label string, one character per array axis in order:
"H"= height (rows),"W"= width (columns),"C"= channels. For example,"HWC"is height × width × channels (NumPy/OpenCV layout);"CHW"is channels × height × width (PyTorch layout). Seepyvisim.typing.value_range (tuple[float, float]) – The
(low, high)range the input values live in; converted into the canonical[0, 255]range.
- Returns:
Feature descriptors (NumPy array).
- Return type:
- detect_and_extract(image)[source]¶
Detect the keypoints and extract their descriptors.
Parameters¶
- image2D array
Input image.
- extract(image)[source]¶
Extract the descriptors for all keypoints in the image.
Parameters¶
- image2D array
Input image.
- extract_batch(images, /, *, dims='HWC', value_range=(0.0, 255.0))¶
Extracts features from a batch of images.
Returns one
(N_i, D)feature array per image, in input order, since the number of descriptors an image yields varies from image to image. This default implementation extracts one image at a time; extractors that can do the whole batch in one go (e.g. a single forward pass through a neural network) override it.- Parameters:
images (Sequence[_Buffer | _SupportsArray[dtype[Any]] | _NestedSequence[_SupportsArray[dtype[Any]]] | bool | int | float | complex | str | bytes | _NestedSequence[bool | int | float | complex | str | bytes]]) – Batch of images, each a
MatLike(NumPy array, torch tensor or array-like) normalized to a canonicaluint8(H, W, C)image before extraction.dims (str) – Axis-label string, one character per array axis in order:
"H"= height (rows),"W"= width (columns),"C"= channels. It applies to every image of the batch. Seepyvisim.typing.value_range (tuple[float, float]) – The
(low, high)range the input values live in; converted into the canonical[0, 255]range.
- Returns:
One
(N_i, D)feature array per input image.- Return type:
- classmethod from_dict(state, **kwargs)¶
Rebuilds an object from a state dictionary (see
to_dict()to see the expected format).- Parameters:
state (dict[str, Any]) – A JSON-safe description of the object.
kwargs (Any) – Objects the state cannot describe, forwarded by
load_from_disk(). Implementations that accept none raise an error ifkwargsis not empty.
- Returns:
A ready-to-use instance.
- Return type:
FeatureExtractorBase
- classmethod load_from_disk(path, **kwargs)¶
Loads an object previously saved with
save_to_disk().Not every part of an object survives serialization: an arbitrary callable such as a torchvision transform has no portable description, so it is left out of the file. Pass such an object back here as a keyword argument.
- Parameters:
kwargs (Any) – Objects the file cannot hold, forwarded to
from_dict().
- Returns:
A ready-to-use instance.
- Raises:
FileNotFoundError – If
pathdoes not exist.ValueError – If the file is not a valid file of this kind or was saved by a different class.
TypeError – If the class does not take one of
kwargs.
- Return type:
_SerializableT
- save_to_disk(path)¶
Saves the serialized state of this object to a file.
- to_dict()¶
Serializes this object into a JSON-safe state dictionary.
The mapping holds the output of
_state()plus the format version under"format_version"and the class name under"__class__". Arrays may be embedded as__ndarray__nodes, which the serialization layer stores as binary tensors.- Returns:
A JSON-safe description suitable for
from_dict().- Return type:
- property deltas¶
The sampling distances of all octaves