Two similar objects cross. The tracker detects both in every frame but still has to decide which new detection belongs to which earlier path. That choice depends on more than the current image: it rests on assumptions about motion, appearance, and uncertainty. Before deep-learning trackers became common, researchers made those assumptions explicit and tested how they held up as video conditions changed.
Tracking before object detection became routine
Early visual tracking covered several related problems. Some systems followed a single marked point; others estimated the outline of a moving region or maintained multiple object trajectories. A tracker might be initialized by hand, receive candidate locations from a separate detector, or find motion against a relatively stable background. Those starting conditions mattered. Following a known object through a short clip was easier than finding every object and preserving its identity throughout a long recording.
Computing limits shaped the methods, too. A camera delivered the next frame before an expensive algorithm could search every possible location and shape. Trackers predicted where to look, limited the search to a neighborhood, and described objects with compact measurements such as position, velocity, color histograms, or edge patterns.
From local motion to an object path
Optical flow and point tracks
Optical flow estimates apparent movement between frames. One common assumption is brightness constancy: a small patch of the same surface keeps roughly the same intensity as it moves. Local methods, notably Lucas–Kanade tracking, use image gradients within a window to estimate displacement. A distinctively textured patch can often be located efficiently in the next frame.
But a flat patch offers little evidence of direction, and motion blur or changing illumination can undermine brightness constancy. An edge may show movement across it but not along it—the aperture problem. Researchers therefore favored corners and other well-textured features. They also often tested a track by moving it forward to the next frame and back again to see whether it returned to a plausible starting point.
Point motion is not necessarily object motion. Points on a walking person move differently as limbs articulate, while background features may fall within the same search window. Grouping points that moved coherently was one way to build an object-level trajectory from local measurements.

Foreground regions and contours
For a stationary camera, background subtraction offered another starting point. The system compared each frame with a model of the usual scene and marked sufficiently different pixels as candidate foreground. Simple frame differencing reacted quickly but could leave holes inside moving objects. Adaptive background models handled gradual changes better, although shadows, reflections, and swaying vegetation could still be mistaken for foreground.
Connected foreground regions could then supply bounding boxes or centroids for tracking. That worked when objects were separate. If two people overlapped, their regions could merge into one component and split again later. Contour-based methods instead moved an outline toward image boundaries using edge evidence and shape constraints. They captured more geometric detail but were sensitive to their initial outline and to unclear boundaries.
Predicting where the object should appear
A tracker needs a temporal model because measurements are noisy and sometimes missing. The Kalman filter became a standard choice when position and velocity could be approximated by linear dynamics with Gaussian uncertainty. It first projects the state into the next frame, then combines that prediction with a new observation. An uncertain observation carries less weight. The covariance matrix records how uncertain the estimate is, rather than treating the predicted location as exact.
That approach suits smooth motion and short gaps, but abrupt turns or competing explanations are harder to handle. Particle filters offered more flexibility by representing possible states as weighted samples, keeping several plausible locations in play. At each frame, the tracker propagated the particles, scored them against the image, and resampled in favor of stronger candidates. It still needed enough particles to cover likely motion—and an appearance measure that could tell the target from clutter.
Mean shift tracking took a different approach, moving a search window toward the best match for a target’s appearance, often represented by a color histogram. It was computationally attractive, but a similarly colored background could pull the window off target. Updating the appearance model continuously introduced another risk: after a partial occlusion, the tracker might start learning the occluder instead.
What the measurements actually told researchers
An early tracker’s output makes more sense when its observation model is clear. Three common sources of evidence had different weaknesses:
- Position or motion: narrowed the search, but offered little help when objects crossed or suddenly changed direction.
- Color and texture: helped distinguish a target from its surroundings, but were vulnerable to illumination changes and look-alike objects.
- Edges and shape: provided evidence for distinctive outlines, but became less reliable when boundaries were hidden or the object deformed.
Combining cues could make a tracker less prone to failure, as long as it did not count correlated or unreliable measurements as independent evidence. A color match might be persuasive in one sequence and nearly useless in another. Academic comparisons therefore needed to say what information a method received at initialization and whether it updated that information while tracking.

Occlusion, identity, and the test sequence
Occlusion shows why predicting a location is different from confirming an identity. If someone walks behind a pillar, a motion model can carry the estimated path through the gap. When the person reappears, the tracker must decide whether the new observation belongs to that path. In scenes with several objects, researchers matched observations to tracks using proximity, motion consistency, and appearance similarity. Assigning each observation to the nearest track was simple, but it could swap identities at a crossing. More elaborate methods considered multiple assignments or waited for later frames before settling on one.
Evaluation had to reflect those differences. A single-object experiment might report the error in the predicted center location or the overlap between predicted and annotated bounding boxes. A multi-object experiment also needed to count missed objects, false tracks, and identity switches. A method could place boxes accurately while repeatedly swapping identities, or keep identities intact while its boxes drifted. Comparisons were more informative when test sequences varied camera motion, crowding, object scale, and occlusion length instead of relying on one average.
Suppose two pedestrians overlap for six frames. Before the overlap, each track has a predicted position and an appearance description. During it, the tracker may see only one merged region, so uncertainty grows. When the pedestrians separate, assigning each region to the nearest predicted position may reverse their identities if they crossed behind the obstruction. The experimental record should count that reversal as an identity switch, not merely a temporary increase in location error.
