Info / Build a try-on

Build a simple virtual try-on

A get-started guide to a glasses try-on that runs in the browser: the webcam, face tracking, a pair of glasses that follows the head, and the basics that make it look right. It takes an afternoon by hand, and less with a coding assistant.

Browser try-on demo: a woman on a video call wearing the sample black glasses, placed by the starter code
The try-on starter demo placing the sample glasses on a woman looking over her shoulder
The try-on starter demo placing the sample glasses on a man facing the camera

Try the finished demo · Download the starter files (zip) · Sample frame (.glb)

What you’ll build

A web page that asks permission, starts the camera, finds the face and places a 3D pair of glasses on it. Everything runs on the device and nothing is uploaded. It shows how a style looks on someone, not whether it’s their size.

What you need

  • MediaPipe Face Landmarker to find the face and how the head is turned. Free, Apache 2.0.
  • three.js to draw the glasses. Free, MIT. Babylon.js works too; switch it to a right-handed scene so MediaPipe’s numbers can be used as they are.
  • MediaPipe’s canonical face model, to hide the parts of the glasses that are behind the head.
  • A studio HDRI for lighting, for example from Poly Haven (free, CC0).
  • A glasses model as a .glb. Start with our sample frame.
  • A page served over https, or from localhost while you build. Browsers only allow the camera there.
The sample frame: plain black glasses with clear lenses
The sample frame.

The starter files put it all together. Here is the top of main.js, with the libraries loaded from a CDN:

import * as THREE from 'three';
import { GLTFLoader } from 'three/addons/loaders/GLTFLoader.js';
import { OBJLoader } from 'three/addons/loaders/OBJLoader.js';
import { HDRLoader } from 'three/addons/loaders/HDRLoader.js';

const GLASSES_URL = 'sample-glasses.glb';
const FACE_URL = 'canonical_face_model.obj';   // MediaPipe's canonical face, in cm
const ENV_URL = 'studio_small_08_1k.hdr';      // any studio HDRI from polyhaven.com
const VISION = 'https://cdn.jsdelivr.net/npm/@mediapipe/tasks-vision@1.0.1';
const MODEL = 'https://storage.googleapis.com/mediapipe-models/face_landmarker/'
  + 'face_landmarker/float16/1/face_landmarker.task';
const FIT = 0.9;      // frame width as a share of the face's width
const SMOOTH = 0.5;   // 1 = no smoothing; lower = steadier but laggier
const DROP = 0.6;     // cm to move the glasses down the face (see step 6)

1. Ask before using the camera

A try-on processes images of people’s faces, so ask first, keep everything on the device and give people a clear way to say no. Don’t load the face model until they agree. Biometric privacy laws, such as Illinois’ BIPA and the EU’s GDPR, apply to try-ons, so check what applies where you operate.

<style>
#stage { position: relative; width: min(100vw, 960px); margin: 0 auto; transform: scaleX(-1); }
#stage video { display: block; width: 100%; height: auto; }
#stage canvas { position: absolute; inset: 0; width: 100%; height: 100%; }
</style>

<div id="consent">
  <p>This try-on uses your camera to follow your face. Everything runs on this device:
    no image or face data is uploaded or stored, and it stops when you close the page.</p>
  <button id="yes">Start the camera</button>
  <button id="no">No thanks</button>
</div>
<div id="stage" hidden><video id="video" playsinline muted></video></div>
// 1. Ask first. Nothing starts, and the face model is not downloaded, until "yes".
document.getElementById('no').onclick = () => {
  document.getElementById('consent').textContent = 'Try-on is off. Nothing was started.';
};
document.getElementById('yes').onclick = start;

2. Start the camera

Show the video and lay a transparent canvas exactly over it. The CSS above mirrors the video and the canvas together, so it behaves like a mirror and the tracking still lines up.

// 2. Start the camera
video.srcObject = await navigator.mediaDevices.getUserMedia({
  video: { facingMode: 'user', width: 1280, height: 720 }, audio: false,
});
await video.play();

You should see yourself, mirrored.

3. Track the face

Ask the Face Landmarker for the facial transformation matrix: where the head is and how it is turned, relative to MediaPipe’s canonical face. If the graphics card can’t be used, it falls back to the CPU.

// 3. Track the face (MediaPipe loads only now, after consent)
const { FilesetResolver, FaceLandmarker } = await import(VISION + '/vision_bundle.mjs');
const files = await FilesetResolver.forVisionTasks(VISION + '/wasm');
const options = (delegate) => ({
  baseOptions: { modelAssetPath: MODEL, delegate },
  runningMode: 'VIDEO',
  numFaces: 1,
  outputFacialTransformationMatrixes: true,
});
const landmarker = await FaceLandmarker.createFromOptions(files, options('GPU'))
  .catch(() => FaceLandmarker.createFromOptions(files, options('CPU')));

4. Match MediaPipe’s camera

MediaPipe calculates that matrix for a camera at the origin with a 63° vertical field of view, measuring in centimeters. Give the 3D camera the same settings and the video’s aspect ratio, and anything placed under the matrix lands on the face.

// 4. A 3D scene the size of the video, with the camera MediaPipe assumes:
//    63 degrees vertical field of view, at the origin, units in centimeters.
const w = video.videoWidth, h = video.videoHeight;
const renderer = new THREE.WebGLRenderer({ alpha: true, antialias: true });
renderer.setSize(w, h, false);
document.getElementById('stage').appendChild(renderer.domElement);
const scene = new THREE.Scene();
const camera = new THREE.PerspectiveCamera(63, w / h, 1, 10000);
const anchor = new THREE.Group();       // follows the head; glasses go inside it
anchor.matrixAutoUpdate = false;
scene.add(anchor);

5. Hide what’s behind the head

Load the canonical face and draw it into the depth buffer only. It stays invisible, but it hides the temples where they pass behind the head.

// 5. Hide what's behind the head: MediaPipe's canonical face, drawn into depth only
const faceObj = await (await fetch(FACE_URL)).text();
const occluder = new OBJLoader().parse(faceObj);
occluder.traverse((m) => {
  if (m.isMesh) m.material = new THREE.MeshBasicMaterial({ colorWrite: false });
});
occluder.renderOrder = -1;
anchor.add(occluder);
Glasses on a turned head without the occluder: the far temple shows across the face
Without the occluder
The same frame with the occluder: the far temple is hidden behind the head
With the occluder

Turn your head: the far temple should disappear behind it.

6. Put the glasses on

Size the frame to the face and put the lens centers a little below eye height, just in front of the bridge of the nose. The positions come straight from the canonical face, whose vertices are in landmark order: 168 is the bridge of the nose, 33 and 133 are the corners of one eye, and 127 and 356 are the sides of the face.

MediaPipe canonical face with the sample glasses placed on it; dots mark landmarks 168 (red), 33 and 133 (blue), and 127 and 356 (green)
The canonical face: bridge of the nose (red), eye corners (blue), sides of the face (green).
// 6. Put the glasses on: sized to the face, lens centers a little below eye
//    height, just in front of the bridge of the nose. The OBJ's vertices are
//    in landmark order, so v[168] is landmark 168.
const v = faceObj.split('\n').filter((l) => l.startsWith('v '))
  .map((l) => new THREE.Vector3(...l.trim().split(/\s+/).slice(1).map(Number)));
const faceWidth = v[356].x - v[127].x;
const eyeY = (v[33].y + v[133].y) / 2;
const glasses = (await new GLTFLoader().loadAsync(GLASSES_URL)).scene;
const size = new THREE.Box3().setFromObject(glasses).getSize(new THREE.Vector3());
glasses.scale.setScalar(FIT * faceWidth / size.x);
glasses.position.set(0, eyeY - DROP, v[168].z + 0.5);
anchor.add(glasses);

Nudge it down. Placed exactly at eye height, the glasses look too high, for two reasons. Real frames sit with the eyes a little above the middle of the lenses. And MediaPipe’s head position is worked out for a typical camera, so on a real webcam it tends to land a little high, more so when the face is near the top of the picture. DROP moves the glasses down the face to cover both; 0.6 cm suited the clips we tested, so adjust it to taste.

This fits the frame to the face, which keeps things simple. It shows how a style looks, not whether it’s the right size.

7. Light it

Metal and glossy acetate need something to reflect. A studio HDRI as the environment is enough to start.

// 7. Light it with a studio environment map
const pmrem = new THREE.PMREMGenerator(renderer);
scene.environment = pmrem.fromEquirectangular(await new HDRLoader().loadAsync(ENV_URL))
  .texture;

8. Follow the head

Run the tracker on every video frame and move the anchor to the result. A little smoothing takes out the jitter, and the glasses hide when no face is found.

// 8. Follow the head every frame, with a little smoothing
const target = new THREE.Matrix4();
const tp = new THREE.Vector3(), tq = new THREE.Quaternion(), ts = new THREE.Vector3();
const p = new THREE.Vector3(), q = new THREE.Quaternion(), s = new THREE.Vector3();
let seen = false;
renderer.setAnimationLoop(() => {
  const res = landmarker.detectForVideo(video, performance.now());
  const found = res.facialTransformationMatrixes.length > 0;
  if (found) {
    target.fromArray(res.facialTransformationMatrixes[0].data).decompose(tp, tq, ts);
    const k = seen ? SMOOTH : 1;
    p.lerp(tp, k); q.slerp(tq, k); s.lerp(ts, k);
    anchor.matrix.compose(p, q, s);
    anchor.matrixWorldNeedsUpdate = true;
  }
  seen = found;
  anchor.visible = found;
  renderer.render(scene, camera);
});

Move around: the glasses should stay on your face.

What a shopping-grade model needs

The code is the easy part. For a try-on people rely on when they buy, the model has to be the frame they will receive.

  • It matches the real frame: shape, proportions, color, finish and logos, checked against the product.
  • The lenses are true to life: tint, gradient and mirror as they really are, and see-through where they should be.
  • Real size, in meters, in one consistent orientation: +Y up, the front facing +Z, and the origin at the center of the bridge, as the code above expects.
  • Every colorway has its own model, so every option a customer can buy is one they can try.
  • It is light enough for the web: a few megabytes at most, with clean, well-organized parts.

That’s what we do: accurate 3D models of every frame and colorway, built from the real product. Eyewear digitization

From demo to shopping tool

A demo like this is a good start. Making it look real, sit right on every face and run smoothly on every phone takes considerably more work. If you’d rather not build that part yourself, talk to us.

Sources and licenses

  • MediaPipe, including the canonical face model: Apache 2.0.
  • three.js: MIT.
  • Studio Small 08 HDRI from Poly Haven: CC0.
  • Sample frame and starter code: © Convex3D LLC. Free to use for testing and learning.
  • Demo footage: Artem Podrez, ShotPot, Moe Magners and Tima Miroshnichenko on Pexels.

Keep reading

3D models for virtual try-on

Try-on-ready .glb models of every frame, for any provider, AR effect or 3D viewer.

Interactive 3D viewer

Turn, tilt and try on live 3D models in the browser.

Eyewear digitization

How we digitize whole catalogs, and keep them current.

FAQ

Scanning, AI, ownership, formats, CAD, try-on and pricing, answered briefly.