hello@bytesgenx.com
All guides
Computer Vision · See the service
// Guide · Computer Vision · 9 Sept 2026

Automating warehouse inventory counts with computer vision

What actually works when you replace manual cycle counts with cameras: where to mount them, which model to reach for, and the failure modes nobody warns you about.

DisciplineComputer Vision
Read7 min
Published9 Sept 2026
Written forOps & engineering leads

The short version

  • 01Decide whether you are counting labels or objects before anything else — it changes the cost by an order of magnitude.
  • 02Chokepoint cameras are the cheapest proof; equipment-mounted cameras buy coverage; drones buy the top racks.
  • 03Small purpose-trained models beat large general ones here, and they run on-site.
  • 04Occlusion and lighting are placement problems, not modelling problems.
  • 05The hard question is what happens when the camera and the WMS disagree.

Manual cycle counting is the tax every warehouse pays. Someone walks the aisles with a scanner, reads labels off pallets, and types numbers into a WMS that was already wrong. The count takes days, it interrupts picking, and by the time it finishes the floor has moved on.

Computer vision can take most of that work away. It also fails in ways that are specific, predictable, and almost never discussed in vendor demos. This is a practical account of both.

Decide what you're actually counting

The single biggest driver of cost and accuracy is a question most teams skip: are you counting labels or objects?

Label reading means detecting a barcode, QR code, or printed SKU on each pallet or carton and decoding it. You get an exact identity — this is pallet LP0049812, containing SKU 4471. Accuracy is effectively a solved problem when the label is visible and in focus, because barcode symbologies carry their own error correction. Your engineering problem is entirely about optics and access: can a camera see the label, at sufficient resolution, often enough?

Object counting means detecting the things themselves — cartons on a pallet, pallets in a bay, items on a shelf — without reading any identity. You get a count and a location but not a SKU. This is a genuine detection problem, and its accuracy depends on your stock's visual variety, occlusion, and lighting.

Most warehouses need both, but they need them in different places. Label reading handles the high-value, identity-critical work: what is in this bay. Object counting handles occupancy and anomaly detection: this bay should have four pallets and has three.

Choose where the camera lives

There are three viable mounting strategies and they are not interchangeable.

Fixed cameras at chokepoints

Cameras at dock doors, wrapping stations, and aisle entrances see inventory as it moves. This is the cheapest and most reliable option, because you are imaging goods at a known distance, under controlled lighting, usually at a predictable orientation.

The limitation is that you only observe transitions. You know what entered and left a zone, not what is currently sitting in it. If your WMS drifts, a chokepoint system inherits the drift — it tracks deltas, not ground truth.

Cameras on the equipment

Mounting on forklifts and reach trucks turns your existing traffic into a survey fleet. Every time an operator drives an aisle, you image it. Coverage follows picking activity for free, and no new vehicle has to be justified.

The trade-off is control. You get whatever angle, speed, and motion blur the operator's route produces. Pose estimation becomes essential — an image is worthless if you cannot say which bay it shows — and that usually means fusing wheel odometry, a fiducial marker scheme, or UWB anchors with the visual data.

Autonomous drones

A drone flying a fixed nightly route gives you what neither of the others does: complete, scheduled, ground-truth coverage of every bay including the top racks a person cannot easily reach. For high-bay facilities this is often the only way to see the upper levels at all.

It is also the most involved option. Indoor flight means no GPS, so navigation runs on visual-inertial odometry, fiducials, or LiDAR. You need charging infrastructure, a flight-safety story, and a plan for what happens when the aisle is not empty.

Pick the smallest model that works

There is a strong pull toward large, general-purpose vision models. Resist it for this problem. Warehouse inventory is a narrow, repetitive visual domain, which is exactly where small purpose-trained models beat general ones on every axis that matters — latency, cost, and the ability to run on-site.

TaskReach forWhy
Barcode / label decodeA dedicated decoder librarySymbologies carry error correction; this is not an ML problem
Pallet & carton detectionA small object detector (YOLO-class)Few classes, high repetition, runs on edge hardware
Bay occupancyDetection plus a geometric bay mapThe map does the reasoning; the model only finds objects
Damage / anomalyAnomaly detection on normal stockYou have thousands of normal examples and almost no damaged ones
Text on unlabeled stockOCRHandles printed carton markings barcodes miss

The pattern across all of these: keep the model narrow and put the intelligence in the surrounding geometry. A detector that finds pallets, combined with a map of where bays physically are, gives you bay-level occupancy. Trying to train a single model to output "bay A-14-3 contains 3 pallets" end-to-end is far harder and far more brittle.

Run inference where the cameras are

Warehouse networks are frequently the least-loved infrastructure in the building. Streaming continuous video from dozens of cameras to a cloud GPU means paying for bandwidth you may not have, and accepting that the system stops working during any connectivity blip.

On-device inference inverts that. Each camera or vehicle runs its own detector and sends structured results — a few hundred bytes of "bay, timestamp, count, confidence" — rather than a video stream. The uplink requirement collapses, the system degrades gracefully offline, and no footage of your operation leaves the site, which materially shortens the security review.

Reserve the cloud for what it is genuinely better at: aggregation, historical trends, retraining pipelines, and dashboards.

3
Mounting options
Chokepoint, equipment-mounted, and autonomous — each answers a different question.
1zone
Sensible pilot
Instrument one representative aisle, not the building.
2systems
Reconciliation
Vision and the WMS will disagree. Decide in advance which one wins.

The failure modes nobody demos

These are the things that turn a working pilot into a stalled rollout.

Occlusion is the whole game. A pallet behind another pallet is invisible, and no model solves that. This is a camera-placement problem, not a modelling problem. Budget for multiple viewpoints per aisle, and accept that some positions are simply unobservable and must be inferred from movement history rather than seen.

Lighting varies more than you expect. Warehouses mix daylight from high windows, sodium or LED fixtures, and deep shadow between racks. The same bay looks materially different at 9am and 9pm. Training data collected over a single shift will not survive contact with the other ones.

Stock appearance drifts constantly. Packaging changes, seasonal SKUs arrive, shrink-wrap colour changes, a supplier switches carton design. Every one of these shifts your input distribution. A model trained once and never revisited degrades quietly — the counts stay plausible while becoming steadily less true, which is worse than obvious failure.

Reconciliation is a product problem, not an ML one. The hard question is not "what did the camera see" but "what do we do when the camera and the WMS disagree." Does vision win? Does it raise a task for a human? Does it depend on value or confidence? Answer this before you build, because the answer determines your entire alerting and workflow design — and getting it wrong is how a technically successful system ends up ignored.

What a sensible first phase looks like

Instrument one zone, not the building. Pick an aisle that is representative but not business-critical, and run the vision system in parallel with your existing count for a full cycle. You are not trying to replace anything yet — you are building a labelled record of every case where the two disagree, and finding out which side was right.

That disagreement log is the most valuable artefact of the whole project. It tells you your true accuracy, it becomes your training and evaluation set, and it is the evidence that makes the case for a wider rollout to people who are reasonably sceptical of replacing a process they can currently watch a human perform.

Expect the first pass to surface more problems with your existing data than with the cameras. That is normal, and it is usually where the immediate return comes from.