How ValEye Counts Pills

ValEye counts pills from a single overhead photo, on your device. The vision pipeline runs in your browser via OpenCV compiled to WebAssembly and works offline once installed. A typical photo takes 3–5 seconds; a dense or difficult one can take 30 seconds or more, and counting runs in a background thread so the app stays responsive while it works. Every photo you count is also uploaded to ValEye’s own server, where it becomes a test case: your corrections (the +/− nudge, the thumbs-up, the “Wrong” report) are the ground truth the counter is measured against. All of your devices share one history.

The problem sounds easy — "count the blobs" — but real trays have touching pills, shadows, white pills on white counters, engraved trays, and piles. The pipeline below is mainly classical computer vision, tuned against a test corpus of real pharmacy-style photos and synthetic stress tests with exact known counts. On photos where the mask is shredded by glare, a small on-device segmentation model (MobileSAM) runs as a second opinion where the hardware supports it — it never counts on its own; its mask has to win a coverage test against the classical one before it is used.

The pipeline

Loading the vision engine to render the stage demos below — these are computed live in your browser, right now, by the same code the app uses…

1

Estimate the background

The border of the photo is assumed to be tray/counter. Its median color becomes the background model — no fixed color or special tray required.

2

Color-distance segmentation with shadow damping

Every pixel gets a distance-from-background score. The luminance component is down-weighted only when a pixel is darker than the background at the same hue — the signature of a cast shadow. White pills on a white counter keep full weight, so they still register. Otsu's method picks the foreground threshold automatically.

3

Second-mode rescue

If the tray holds both strongly-colored and faint pills, Otsu splits "colored vs everything" and drops the faint ones. We re-run Otsu on just the leftover background; anything it reveals is admitted only if it is a compact, pill-shaped piece (its area must be explainable by its thickness — a disk, not a smudge).

4

Surface refinement

A blob covering more than ~12% of the frame is a plate or tray, not a pill. It gets re-segmented against its own dominant color so the pills sitting on it pop out. An accept-test guards the other direction: if re-segmenting yields crevice fragments instead of several pill-like pieces (i.e. the blob was actually a pile of same-colored pills), the refinement is reverted.

5

Distance transform & markers

The distance transform gives each foreground pixel its distance to the nearest background pixel — pill centers become peaks, and a blob's peak is its half-width. Pill-scale blobs seed one marker per peak cluster (60% of the blob's own peak); blobs thicker than a pill (piles) instead use local maxima, spaced by the estimated pill radius, so each buried pill center still gets a seed.

6

Watershed

The watershed algorithm floods outward from the markers and draws ridge lines where floods meet — which is exactly the boundary between two touching pills. This is the classic separator for touching convex objects.

7

Filter, split, number

Regions are filtered by an absolute size floor (relative to the image), a relative floor (vs the median region), and a minimum thickness — killing specks, rims, and engraving lines. A region much larger than the median (a merged cluster the watershed missed) contributes round(area / median) pills and is badged as a range, e.g. 12–14. Every counted pill gets a numbered badge in the overlay.

Pixel mass: the counting rule that won

A pharmacy count is one medication, so every pill in frame has the same area in pixels. That turns counting into arithmetic: after segmentation, each blob's pixel mass should be an integer multiple of one pill's mass — a lone pill is 1.0×, two touching pills are ~2.0×, a chain of five is ~5.0×. The unit mass is estimated iteratively (start at the median blob, divide each blob by its rounded multiple, re-take the median), and the count is Σ round(blobArea / unitArea).

Three approaches were A/B-tested head to head. Pixel mass won and is the default; the watershed baseline is kept because it fails differently, which is what makes it useful as the second opinion described below.

Watershed still runs — it draws the pill boundaries and badge positions for the overlay — but the number comes from pixel mass. The failed sized experiment taught us why: dividing watershed regions by area amplifies marker noise, while dividing raw blobs (whose outlines come straight from segmentation) is stable.

Current results

Measured on every corpus board, decoded the way a real browser decodes it (regenerated on each release — see the test bench for the live table):

CorpusExactWithin 1
Real photos (the honest number) 43/51  84%46/51  90%
All boards incl. synthetic stress tests 258/266  97%261/266  98%

Those two numbers differ for a reason worth stating: the synthetic boards are rendered with exact known counts and the counter gets 215 of 215 right, so a single combined figure flatters it. The real-photo number is what a user should expect.

ScenarioResult
Scattered same-size pills, overhead, plain matte surface (the intended use) exact, and the great majority of corpus photos are this
Pills touching in a single layerexact on most boards, including 90-pill fused rafts
White pills on a white counterhandled by shadow-aware colour distance
Glossy/specular pillsthe hardest real class — usually exact, off by 1–2 on two corpus boards
Busy or textured surface (wood grain, fabric, a patterned tray) can invent pills — the app refuses these rather than showing a number
Two medications, mixed sizes, or pills stacked on each other out of scope — one medication, one layer
Angled shots, piles with buried pills, pills cut off by the frame out of scope — shoot straight down and pull back so nothing touches the edge

When ValEye is unsure — and how much to trust it

Every count carries a confidence score, and that score has been calibrated: binned against known-correct answers to check that “80%” really means right 80% of the time. It mostly errs on the safe side (the 0.8–0.9 band is right 94% of the time), and below 0.65 the app declines to show a number at all and asks for a retake.

Three independent warnings run on top of the count:

Between them, every wrong count in the current corpus raises at least one warning. That is the design goal: be right, and when you are not, say so.

ValEye is a second pair of eyes, not a dispensing device. It has no regulatory clearance and no clinical validation. For anything that matters, count it yourself and use ValEye to check your count — not the other way round.

Test methodology

The test bench runs the full pipeline in your browser over the corpus: real CC-licensed photos (with hand-counted expectations) plus synthetic stress tests rendered with exact known counts — touching chains, tiny/large pills, gradients, sensor noise, motion blur, vignettes, white-on-white. Grading: pass within max(1, 10%), warn within 30%, fail beyond. The same suite runs headless via node tools/count-cli.mjs testdata --ab for algorithm A/B comparisons.

Stack

Static PWA (no build step, no server): vanilla JS + OpenCV.js 4.9 (single-file WASM, ~10 MB, cached by a service worker for offline use). The counting module is environment-agnostic and runs identically in Node for headless regression testing.