Pedestrian signals,
made audible.

Worn on the chest, eye reads Taiwan's pedestrian signals aloud for blind and low-vision people. When it is unsure, it says it cannot confirm. It never guesses green.

A CanMV K230 research prototype, not yet field-tested. It does not replace a white cane, a guide dog or help from others.

Device voice (translated)

Voice prompts

The device speaks a small, fixed set of short prompts in Mandarin, each with a single clear meaning, so nothing is misheard at a noisy intersection. Below are English translations, each followed by the original prompt the device plays.

Red light. Please wait.
紅燈,請等待The pedestrian signal ahead is red.
Green light. You may cross.
綠燈,可以通行The signal ahead is green, and confidence clears the stricter threshold set for green.
Green ending. Do not start crossing.
綠燈快結束,請勿開始穿越Flashing green detected: not enough time is left to cross safely.
Signal unconfirmed. Do not start crossing.
號誌無法確認,請勿開始穿越A signal previously read as green can no longer be confirmed, and that reading should no longer be relied on.
Signal unclear. Please wait.
號誌不清楚,請稍等A signal is in view, but its color cannot be read reliably.
No signal detected. Turn slowly left and right.
未偵測到號誌,請慢慢左右轉動After a few seconds without a signal in view, the user is asked to turn slowly on the spot to find one.

Filled dot: a steady light Ring: a flashing light Gray dot: a status prompt, not a light

There are also status prompts such as “System started,” “Signal found” and “System error.”

How it works: from image to speech

Every frame passes through five steps. The first two locate the signal; the last three decide whether to speak, and what to say.

1, 2Limit the view to the direction of travel and locate the signal

Green
0.94
Red
0.03
Other
0.03

3Crop the image and classify the color (values are illustrative)

Several frames in a row read green

4, 5Speak only once the frames agree

How a single frame is processed (illustration). The board handles about 14 frames per second.
  1. Limit the view

    Only the central region of the image, in the user's direction of travel, is analyzed, so signals at neighboring crossings are not misread.

  2. Locate the signal

    A detection model (YOLO26n) draws a box around the pedestrian signal within that region.

  3. Read the color

    A 96×96 crop is cut from the full-resolution image, and a color classifier (YOLO11n-cls) labels it red, green or other. Clear red light inside the box overrules any green reading.

  4. Check over time

    The state changes only when several consecutive frames agree, so a single misread frame never turns into speech.

  5. Speak

    The K230 board plays the prompt itself. No phone or network connection is needed.

Vision: a two-model design

Locating the signal and reading its color are handled by two lightweight models. Locating needs a wide view; reading the color needs only a small patch, which can be cropped straight from the full-resolution image for better reads of distant signals.

Detection model

Task
Locate the pedestrian signal
Architecture
YOLO26n
Input
640 × 640 travel-direction region
Output
Pedestrian-signal bounding boxes
On-device inference
55.9 ms

Color classifier

Task
Red, green or other
Architecture
YOLO11n-cls
Input
96 × 96 signal crop
Output
3-class probabilities
On-device inference
3.2 ms

Both models are converted to int8 kmodels with nncase and run on the K230's KPU (neural-network accelerator), fully offline. Training used GPU time provided by Kaggle.

Collecting and cleaning the training data

  • 2,455

    photos of Taiwanese intersections (our own dataset), 1,567 of them with pedestrian-signal labels.

  • 1,364

    kept after OpenCLIP removed photos taken from scooters and vehicles, leaving only the pedestrian's point of view.

  • 630

    vehicle signal heads added as negative examples, cutting vehicle signals misread as a pedestrian green from 12 to 0.

  • 80%

    is the bar for any new data source: a manual spot check of its labels must reach 80% accuracy, or the whole batch is rejected. A second batch of Wikimedia Commons images scored 40.6% and was not used.

  • 19

    images that no model has ever seen are held back as a locked test set. It is opened only at release, to confirm that scores are not inflated by overlapping data.

The data also includes photos from the board's own camera, close-range synthetic images, and pedestrian-signal data from other countries whose labels were re-checked one by one.

Training data counts

Datasets behind the current candidate models. The training set is used for learning, the validation set for choosing the best epoch, and the test set only for the final evaluation.

Training data counts for the two models
ModelTrainingValidationTestTraining set makeup
Detection model roi_v4 2,582396125 1,043 real Taiwanese intersection photos, 1,500 close-range synthetic images, 39 indoor negatives; 2,648 labeled boxes in total
Color classifier v12_hf 6,474763250 96 × 96 crops: green 2,078, red 2,196, other 2,200

Better to say “cannot confirm” than to wrongly say “green.”

  • A zero-tolerance error

    Reporting a red light as green could send the user into traffic. No model may be deployed to the board unless this error count is 0 in offline evaluation.

  • Asymmetric confidence thresholds

    The color classifier must be at least 0.8 confident before green is announced; red needs 0.6. The two mistakes have different consequences, so their thresholds differ too.

  • Never default to green

    When the image is unclear, the signal is too small, or consecutive frames disagree, the system reports that it cannot confirm the signal. New features may only make the system more cautious, never more likely to announce green.

  • When a green is lost

    If a signal read as green can no longer be confirmed, the system says “do not start crossing” rather than “stop”: the user may already be in the crosswalk, and stopping suddenly could be more dangerous.

Results

Offline evaluation of the current candidate pair (detection model roi_v4 + color classifier v12_hf). The first measure of this project is never reporting a red light as green; overall accuracy comes second.

Red or non-signals reported as green
0every frame of the test and validation sets
Vehicle signals misread as a pedestrian green
0/28vehicle signals from the Australian NSW public dataset
On-device speed
14.4frames per second (measured on the K230)

Reading accuracy

Dots show accuracy; lines show the 95% confidence interval (Wilson). The fewer the samples, the wider the interval.

Test setwithin the device's view 91.9%
Validation setwithin the device's view 87.8%
Unseen by the detector32 images 90.6%
Board camera photos44 images 86.4%
“Within the device's view” excludes portrait close-ups: the board's camera is landscape, so it never produces such frames in use. With them included, test and validation accuracy are 83.9% and 82.0%.

Why 57 frames were misread

Out of 179 frames containing a signal in the test and validation sets.

  • Actual misreads
  • Skipped by rule
Portrait close-ups 28
Small, distant signals 17
Color reading 11
Wrong signal chosen 1
All 4 swaps were green read as red, an error in the safe direction. Skipping close-ups is deliberate: removing the rule would raise accuracy but add 1 dangerous frame.

Vehicle signals misread as a pedestrian green

Versions of the color classifier, each paired with the same detection model.

  • Current candidate
  • Other versions
v4on the board now 7
v11 2
v12no vehicle negatives 12
v12_hfcandidate 0
v12 and v12_hf use the same data; the only difference is the 630 vehicle-signal negatives.
View all data as tables
Offline evaluation of the candidate pair (roi_v4 + v12_hf)
MeasureResultRate95% confidence interval
Test set (within the device's view)102 / 11191.9%85.3% to 95.7%
Validation set (within the device's view)158 / 18087.8%82.2% to 91.8%
Images the detector never trained on29 / 3290.6%75.8% to 96.8%
Photos from the board's camera38 / 4486.4%73.3% to 93.6%
Test and validation sets (portrait close-ups included)n/a83.9%, 82.0%n/a
Red or non-signals reported as green0n/an/a
Vehicle signals misread as a pedestrian green0 / 280%0% to 12.1%
Why frames were misread (57 of the 179 frames containing a signal)
CauseFramesBreakdown
Portrait close-ups (skipped by rule)28Signal within about 2 m
Small, distant signals17Not detected 11, too small to read 5, low confidence 1
Color reading11Red and green swapped 4 (all green read as red), read as other 3, low confidence 2, adjacent-signal guard 2
Wrong signal chosen1n/a
Vehicle signals misread as a pedestrian green (color classifier versions, detection model roi_v4)
Color classifierCount
v4 (on the board now)7
v112
v12 (no vehicle-signal negatives)12
v12_hf (candidate)0

Limits of this evaluation: the detection model's training data overlaps some of the test images, so the 32 images it never trained on are reported separately; with so few samples, their interval is wide. Field tests at real intersections have not been done yet.

Current status: the board currently runs an earlier pair (detection model roi_v4 + color classifier v4). The candidate pair still needs model conversion and on-device validation.

Please note: this device is a research prototype and cannot replace a white cane, a guide dog or help from others. When crossing a road, rely on your own judgment and your surroundings.

Wearable enclosure

A 3D-printed enclosure measuring 90 × 64 × 22.4 mm, worn on a chest harness with a standard GoPro mount. It is designed to be operated without sight.

Exploded view of the enclosure: an orange lid with a two-prong GoPro mount; below it the green circuit board, brass standoffs and a blue base, held together by a screw at each corner.
Lid, circuit board and base are held by standoffs and screws; the lid is fully open above the heat sink.
Close-up of the top edge: the red KEY button plunger sticks out of the enclosure, with a raised rib on each side.
Tactile ribs on both sides of the button make it easy to find by touch and help prevent presses by clothing.
Close-up of the braille on the top edge: three cells reading e, y, e.
Braille ⠑⠽⠑ (eye) on the top edge, oriented for the wearer to read.

How to use

How the prototype is operated today. Parts still in development are marked in each step.

  1. Put it on

    Attach the enclosure to a GoPro chest harness with the camera facing forward. Turn the M5 thumb screw to adjust the camera's tilt.

  2. Connect power and earphones

    Connect a power bank over USB-C; speech comes out of the 3.5 mm headphone jack. Open-ear headphones (such as bone-conduction models) are recommended, so you can still hear traffic and other surrounding sounds.

  3. Start reading mode

    When reading mode starts, the device says “System started.”

    In developmentThe board currently starts in a photo mode used for data collection, and reading mode must be started from a computer. The final version will start straight into reading mode.

  4. Face the crossing

    Stand behind the curb, facing the crosswalk you want to cross. If you hear “No signal detected. Turn slowly left and right,” turn slowly on the spot without stepping sideways. Once you hear “Signal found,” hold that direction.

  5. Act on the prompts

    After “Green light. You may cross,” still check the traffic by ear and with your cane before crossing. If you hear “Green ending” or “Signal unconfirmed,” do not start crossing.

Open source and data

This project builds on many open-source projects and public datasets, each used under its original license and credited. The full list of data sources, tools and licenses is on a separate page.

Training data sources
8

3 of them under the MIT license: ImVisible PTL, Traffic Lights of New York, nsw_traffic_lights

Tools and services
24

6 of them under the MIT license: OpenCLIP, ONNX Runtime, trimesh, PyVista, html-validate, fonttools

License to watch
AGPL-3.0

Ultralytics YOLO: its terms apply before publishing a model or any commercial use

See all data sources and licenses