3D illustration: a man with a white cane walks forward wearing the VisionNav chest unit and an earpiece, while scan rings and LiDAR points spread across the floor around him

It remembers the room. Then it walks you there.

A chest-worn guide for blind and low-vision people. A camera and a LiDAR map the space as you walk. Ask for the chair by the door and it talks you there in clock directions, then guides your hand the last few inches.

Earpiece “Chair at 2 o'clock, 6 feet.”
  • 292object types
  • ~2 sto answer
  • 0internet
  • 5buttons

The idea

A cane tells you something is there. Other devices warn you, or describe what's in front of you. VisionNav remembers the room, walks you to what you asked for, and guides your hand to it.

Live map · LiDAR 10 Hz
What VisionNav keeps in its head: walls from the LiDAR, objects pinned by the camera.

The real thing

See the prototype at work

The full story and live demos, recorded with the real rig: “VisionNav: what if AI could be your eyes?”

The prototype

Built, wired and worn on the chest

LiDAR on top, a camera in front, buttons on both sides and a Raspberry Pi inside.

The real VisionNav prototype from the front: the round LiDAR on top, a camera window, the VisionNav logo plate, and green MODE and red LOOK buttons on its right side; the TALK and HAND/FACE buttons are on the left side, out of view
The prototype from the side: the green MODE and red LOOK buttons with small printed labels, and the strap
Side view · MODE and LOOK
  1. RPLiDAR C1 360° sweep, 10 times a second
  2. 720p camera objects, hands, faces, labels
  3. TALK + HAND/FACE buttons on the left side
  4. MODE + LOOK buttons on the right side
  5. Raspberry Pi 5 + MPU-6050 IMU inside

How you wear it

Illustration of how VisionNav is worn: the chest unit strapped to the front, an earpiece in one ear and a white cane in the right hand
  1. Earpiece one ear, so the other stays open
  2. Chest unit sees, measures, listens
  3. White cane still yours; VisionNav adds to it
01 · Chest

Sense

Raspberry Pi 5, LiDAR, camera, IMU and five buttons. It boots straight into ROS 2 and runs no AI itself.

02 · AI

Think

Eight AI models, the room map and the route planner turn what the rig senses into a safe route. Fully offline, no cloud.

03 · Ear

Speak

Piper's offline voice in one ear. One sentence, the most urgent thing first, then silence.

A day with VisionNav

From a hotel room to a busy crossing

3D illustration: he walks into a hotel room while a scan ring spreads on the floor; the desk, the bed and the window are tagged, and the earpiece says “Saved this place as hotel room.”

A room it has never seen

Tap SENSORS and walk around. The LiDAR traces the walls, the camera names the desk, the bed and the window, and each one is pinned to the map.

Earpiece“Saved this place as hotel room.”
3D illustration: in a living room he follows a dotted route past the coffee table to a chair tagged “chair · 12 ft” beside the door; the earpiece says “Turn right, to 2 o'clock. 12 feet.”

Take me to the chair

Hold TALK and say it like you'd say it to a friend. It finds that chair on its map, plans around the table and re-plans every second.

Earpiece“Turn right, to 2 o'clock. 12 feet.”
3D illustration: he reaches across a table towards a cup tagged “cup · 0.6 m”; a panel reads 2 cm sideways, 4 cm depth, and the earpiece says “Left 2 inches… Stop. The cup is at your hand.”

Breakfast, the last inch

Hand mode tracks your fingers and the cup in 3D and counts you in until they're within 5 cm sideways and 7 cm deep.

Earpiece“Left 2 inches… Stop. The cup is at your hand.”
3D illustration: in a pharmacy aisle he holds up a medicine box tagged “label”; he asks “Read this label” and the earpiece says “Paracetamol, 500 mg. Take one tablet.”

At the pharmacy

Hold LOOK and ask. A small vision-language model reads the box on the laptop, with no cloud and nothing uploaded.

Earpiece“Paracetamol, 500 mg. Take one tablet.”
3D illustration: two people stand ahead of him; one waving is tagged “Kamal”, the other “new”, and the earpiece says “Kamal on your left, someone new on your right.”

Meeting a friend

Say “Remember Kamal” once. Faces are matched on the device in about 25 ms and never leave it.

Earpiece“Kamal on your left, someone new on your right.”
3D illustration: he waits at a zebra crossing; the traffic light is tagged “signal · RED” and an approaching three-wheeler “15 ft”; the earpiece says “Signal is red. Wait. Three-wheeler coming, 15 feet.”

The crossing home

Outdoors it reads the signal and times every vehicle. Under 6 seconds to impact it warns; under 3 seconds it says “Stop”.

Earpiece“Signal is red. Wait. Three-wheeler coming, 15 feet.”
3D illustration: walking along a pavement, a branch ahead is tagged “head height” and a pothole “8 ft”; the earpiece says “Low branch at head height. Duck.” and “Pothole ahead, 8 feet.”

An evening walk

The branch a cane swings right under, the hole just past its tip. One warning about the most urgent thing, then silence.

Earpiece“Low branch at head height. Duck.”

Everything it can do

Eight abilities, five buttons

Flat illustration: a scan ring on the floor of a hotel room with the desk, window and bed marked
01

Remembers the room

Objects stay on the map after you turn away. A moved object disappears once the camera looks again.

“The cup is behind you, 4 feet.”

Flat illustration: a dotted route across a living room to a chair by the door
02

Takes you there

Ask in plain words. It plans around furniture and people and re-plans every second.

“Turn left, to 10 o'clock. 12 feet.”

Flat illustration: his hand with landmark dots reaching a mug, with 2 cm sideways and 4 cm depth
03

The last inch

Tracks your hand and the object in 3D until they meet: within 5 cm sideways, 7 cm deep.

“Left 2 inches… Stop.”

Flat illustration: holding up a medicine box in front of pharmacy shelves
04

Ask anything

Hold LOOK and ask about what the camera sees. Qwen3-VL answers offline in about 2 seconds.

“The label says: paracetamol, 500 mg.”

Flat illustration: two people ahead, one tagged Kamal and the other tagged new
05

Knows your people

Say “Remember Kamal” once. Faces stay on the device and match in about 25 ms.

“Kamal on your left.”

Flat illustration: a low branch tagged head height and a pothole tagged 8 feet on a pavement
06

Watches the street

Holes, kerbs, head-height branches and vehicles. Warns under 6 s to impact, says “Stop” under 3 s.

“Low branch at head height. Duck.”

Flat illustration: at a zebra crossing, the signal is tagged red and a three-wheeler 15 feet away
07

Talks only when it matters

One sentence at a time, most urgent first. Nothing behind you is tracked or spoken.

“Stop. Three-wheeler coming, 15 feet.”

08

Works with no signal

No cloud, no GPS, no account. Nothing is uploaded, and the indoor map is forgotten when the session ends, by design.

“Session ended. Map cleared.”

Watch it think

What the rig sees, second by second

Short animations built from the system's real behaviour: the same clock directions, distances, thresholds and phrases the rig uses.

Inside

Sense, understand, remember, speak

  1. 01

    Sense

    The camera sees, the LiDAR measures, the IMU feels each turn.

  2. 02

    Understand

    Objects are named, given a true distance with LiDAR-calibrated depth, and tracked with a Kalman filter.

  3. 03

    Remember

    Each object is pinned to the room map after 15+ sightings, so a false detection never sticks.

  4. 04

    Speak

    The most urgent thing first, at least 2.5 s apart. A line that's more than 1.5 s late is dropped.

Eight AI models

  • DetectionYOLOE-11s-segOpen vocabulary, TensorRT FP16
  • DepthDepth Anything V2Indoor and outdoor, rescaled by LiDAR
  • Scene Q&AQwen3-VL 2B4-bit via Ollama
  • Speech inWhisper small.enfaster-whisper on CPU
  • Speech outPiperNatural offline voice
  • HandsMediaPipe Hands21 landmarks, placed in 3D
  • FacesYuNet + SFaceDetect, then recognise
  • Text encoderMobileCLIPNew objects by name

Try it

Five buttons. Press them.

Tap, hold (about a second) or double tap, and read what VisionNav would say. Keys 1–5 work too. Start with SENSORS.

  • sensors
  • indoor
  • hand
  • face
  • route

Turn on JavaScript to try the buttons.

Earpiece
    Every button, every gesture
    ButtonTapHoldDouble tap
    SENSORSCamera, LiDAR and IMU onSensors off—
    LOOKDescribe the sceneAsk a questionVision AI off
    MODEIndoor ↔ outdoorClose map and camera, forget the session—
    HANDHand mode on / offSay an object, your hand is guided to itFace mode on (again: face and hand off)
    TALKWhat is around meSay where to go (indoor)Stop navigation

    Results

    Measured on the real rig

    • 2.2 smedian answer to a scene question26 of 35 timed questions under 3 s
    • 26×faster vision AI after optimisation52 s → 2 s, 4-bit model on the GPU
    • 92.5%of voice commands understood120 recorded commands, text in 1.4 s
    • 6%median depth error against LiDAR834 checks from 0.4 to 8 m

    How it compares

    CapabilityWhite caneObstacle wearablesRecognition devicesRobotic guidesVisionNav
    Warns of obstaclesyes, by touchyesnoyesyes
    Says what objects arenopartialyesnoyes
    Remembers objects out of viewnonononoyes
    Guides you to a described objectnononopartialyes
    Guides your hand to the objectnonononoyes
    No beacons or pre-made mapyesyesyesoften notyes
    Fully offlineyesvariespartialyesyes

    See less. Know more.
    Walk on your own.