ValEye counts pills from a single overhead photo, on your device. The vision pipeline runs in your browser via OpenCV compiled to WebAssembly and works offline once installed. A typical photo takes 3–5 seconds; a dense or difficult one can take 30 seconds or more, and counting runs in a background thread so the app stays responsive while it works. Every photo you count is also uploaded to ValEye’s own server, where it becomes a test case: your corrections (the +/− nudge, the thumbs-up, the “Wrong” report) are the ground truth the counter is measured against. All of your devices share one history.
The problem sounds easy — "count the blobs" — but real trays have touching pills, shadows, white pills on white counters, engraved trays, and piles. The pipeline below is mainly classical computer vision, tuned against a test corpus of real pharmacy-style photos and synthetic stress tests with exact known counts. On photos where the mask is shredded by glare, a small on-device segmentation model (MobileSAM) runs as a second opinion where the hardware supports it — it never counts on its own; its mask has to win a coverage test against the classical one before it is used.
Loading the vision engine to render the stage demos below — these are computed live in your browser, right now, by the same code the app uses…
The border of the photo is assumed to be tray/counter. Its median color becomes the background model — no fixed color or special tray required.
Every pixel gets a distance-from-background score. The luminance component is down-weighted only when a pixel is darker than the background at the same hue — the signature of a cast shadow. White pills on a white counter keep full weight, so they still register. Otsu's method picks the foreground threshold automatically.
If the tray holds both strongly-colored and faint pills, Otsu splits "colored vs everything" and drops the faint ones. We re-run Otsu on just the leftover background; anything it reveals is admitted only if it is a compact, pill-shaped piece (its area must be explainable by its thickness — a disk, not a smudge).
A blob covering more than ~12% of the frame is a plate or tray, not a pill. It gets re-segmented against its own dominant color so the pills sitting on it pop out. An accept-test guards the other direction: if re-segmenting yields crevice fragments instead of several pill-like pieces (i.e. the blob was actually a pile of same-colored pills), the refinement is reverted.
The distance transform gives each foreground pixel its distance to the nearest background pixel — pill centers become peaks, and a blob's peak is its half-width. Pill-scale blobs seed one marker per peak cluster (60% of the blob's own peak); blobs thicker than a pill (piles) instead use local maxima, spaced by the estimated pill radius, so each buried pill center still gets a seed.
The watershed algorithm floods outward from the markers and draws ridge lines where floods meet — which is exactly the boundary between two touching pills. This is the classic separator for touching convex objects.
Regions are filtered by an absolute size floor (relative to the image), a relative
floor (vs the median region), and a minimum thickness — killing specks, rims, and
engraving lines. A region much larger than the median (a merged cluster the watershed
missed) contributes round(area / median) pills and is badged as a range,
e.g. 12–14. Every counted pill gets a numbered badge in the overlay.
A pharmacy count is one medication, so every pill in frame has the same
area in pixels. That turns counting into arithmetic: after segmentation, each
blob's pixel mass should be an integer multiple of one pill's mass — a lone pill is
1.0×, two touching pills are ~2.0×, a chain of five is ~5.0×. The unit mass is
estimated iteratively (start at the median blob, divide each blob by its rounded
multiple, re-take the median), and the count is
Σ round(blobArea / unitArea).
Three approaches were A/B-tested head to head. Pixel mass won and is the default;
the watershed baseline is kept because it fails differently, which
is what makes it useful as the second opinion described below.
Watershed still runs — it draws the pill boundaries and badge positions for the
overlay — but the number comes from pixel mass. The failed sized
experiment taught us why: dividing watershed regions by area amplifies marker
noise, while dividing raw blobs (whose outlines come straight from
segmentation) is stable.
Measured on every corpus board, decoded the way a real browser decodes it (regenerated on each release — see the test bench for the live table):
| Corpus | Exact | Within 1 |
|---|---|---|
| Real photos (the honest number) | 43/51 84% | 46/51 90% |
| All boards incl. synthetic stress tests | 258/266 97% | 261/266 98% |
Those two numbers differ for a reason worth stating: the synthetic boards are rendered with exact known counts and the counter gets 215 of 215 right, so a single combined figure flatters it. The real-photo number is what a user should expect.
| Scenario | Result |
|---|---|
| Scattered same-size pills, overhead, plain matte surface (the intended use) | exact, and the great majority of corpus photos are this |
| Pills touching in a single layer | exact on most boards, including 90-pill fused rafts |
| White pills on a white counter | handled by shadow-aware colour distance |
| Glossy/specular pills | the hardest real class — usually exact, off by 1–2 on two corpus boards |
| Busy or textured surface (wood grain, fabric, a patterned tray) | can invent pills — the app refuses these rather than showing a number |
| Two medications, mixed sizes, or pills stacked on each other | out of scope — one medication, one layer |
| Angled shots, piles with buried pills, pills cut off by the frame | out of scope — shoot straight down and pull back so nothing touches the edge |
Every count carries a confidence score, and that score has been calibrated: binned against known-correct answers to check that “80%” really means right 80% of the time. It mostly errs on the safe side (the 0.8–0.9 band is right 94% of the time), and below 0.65 the app declines to show a number at all and asks for a retake.
Three independent warnings run on top of the count:
Between them, every wrong count in the current corpus raises at least one warning. That is the design goal: be right, and when you are not, say so.
The test bench runs the full
pipeline in your browser over the corpus: real CC-licensed photos (with hand-counted
expectations) plus synthetic stress tests rendered with exact known counts —
touching chains, tiny/large pills, gradients, sensor noise, motion blur, vignettes,
white-on-white. Grading: pass within max(1, 10%), warn
within 30%, fail beyond. The same suite runs headless via
node tools/count-cli.mjs testdata --ab for algorithm A/B comparisons.
Static PWA (no build step, no server): vanilla JS + OpenCV.js 4.9 (single-file WASM, ~10 MB, cached by a service worker for offline use). The counting module is environment-agnostic and runs identically in Node for headless regression testing.