Skip to content

Biometrics · · 3 min read

Face match: embeddings, similarity and thresholds

How a selfie is compared with a document portrait, what the similarity score means, and how to set the pass and review thresholds.

Face match answers a narrow question: is the person in the selfie the person on the document? It does not say whether the document is genuine or whether the selfie came from a live person; those are the document and liveness checks. This guide explains how the comparison works in KYCVerify and how to reason about its score.

From an image to a vector

  1. Detect. YuNet finds faces and five landmarks for each: both eyes, the nose tip and both mouth corners.
  2. Align. The five landmarks are warped onto a fixed template at 112 × 112 pixels, so every face is upright, centred and the same size before it is measured.
  3. Embed. SFace turns the aligned face into an embedding: 128 numbers placed so that photos of the same person land close together.
  4. Compare. The cosine similarity of the two embeddings is the score: 1 means identical direction, around 0 means unrelated.

Both models are openly licensed, from the OpenCV model zoo, and run on the server's CPU through a pure-Rust runtime. No image is sent to a third party to be compared.

Which two faces

The selfie is the centre frame of the liveness step. The reference is the portrait cropped from the document photo; when there is no document portrait, for example in an Aadhaar-only flow, it is the photo inside the UIDAI-signed Secure QR or Offline XML. The check records which one it used.

Reading the score

  • Below 0.30: failed, warning face_mismatch
  • 0.30 to 0.363: review, warning weak_face_match
  • 0.363 and above: passed (SFace's published threshold)
Default thresholds. Between them, a person decides.

SFace's authors publish 0.363 as the cosine threshold for a match, and KYCVerify uses it as the default pass mark. Scores just below it are often the same person in a poor photo, so the default review band sends 0.30 to 0.363 to a reviewer instead of declining.

Tuning for your traffic

  • Raise `review_threshold` to send more borderline pairs to people and fewer to automatic declines.
  • Raise `threshold` to demand stronger matches before approving automatically; expect more reviews.
  • Look at the scores of sessions your reviewers approved and declined. If approved pairs cluster just under the threshold, the band is too narrow.
  • Old document photos, heavy glare and very small portraits lower scores for genuine people. Image-quality checks on the document step catch some of these before they reach face match.

The same faces across sessions

The duplicate check reuses the embeddings. At submit, the selfie is compared with the faces of other approved, in-review, processing and declined sessions of the same app, and a similarity at or above duplicate.face_threshold (0.55 by default, deliberately stricter than the match threshold) flags a possible repeat applicant. Document fingerprints (HMAC-SHA256 of the document number, keyed with a secret derived from the service's master key (kept outside the database) and salted with your organisation id; for Aadhaar, of the last four digits with the name and birth date) catch the same document even with a different face. Either sends the duplicate check to review; neither fails it on its own (a session can still be declined if another check fails with auto-decline on).

What it cannot do

  • It cannot tell a live face from a photo of a face; that is liveness.
  • It cannot tell a genuine document from a well-made fake bearing the applicant's photo; that is why the MRZ, expiry and issuer are checked, and why sensitive flows add a person.
  • Like every face model, accuracy varies with image quality, age gap, pose and lighting. KYCVerify publishes no accuracy figure for its own pipeline; measure it on your traffic.

Try it without a session: POST /v1/checks/face-match with image_a and image_b returns the similarity, the threshold and how many faces each image held.

See it working on your own data.

Every guide describes code you can run in the sandbox today.