
Corvus ISR offers a wide-area motion imagery (WAMI) exploitation product that has taken an important step towards transparent, reproducible benchmarking. Their recent publication of the public tracker benchmark demonstrates a rigorous, fixed-seed evaluation comparing two tracker models on an identical synthetic scene with perfect ground truth. This approach emphasizes the importance of engineering discipline in developing reliable tracking solutions.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The two models under test are distinctly different: v1, the “greedy nearest-neighbour” baseline, implements a simple, two-pass greedy association with constant-velocity prediction and fixed 2s coasting. In contrast, v2 employs a sophisticated “confirmed-track auction” method, utilizing a three-tier auction, velocity-consistency gating, and noise-scaled reservation pricing. Both models share identical detection properties, isolating tracker differences as the variable of interest.
Headline results show that v2 outperforms v1 significantly, with ID switches per minute reducing by over 42% across different scenarios. For example, under a baseline configuration with 150 movers at 2fps, ID switches decreased from 2,042 to 1,183. Similarly, in dense scenes with 400 movers, switches dropped from 14,032 to 8,040, demonstrating the tracker’s robustness under stress. These measurements are derived from perfect ground truth, ensuring they reflect true system behavior, not sensor limitations.
The metrics used are intentionally strict, with the ID switch count counting every change of track identity, including fragmentations and re-acquisitions, exceeding the standard MOT-challenge definition. Publishing such detailed failure numbers underlines Corvus ISR’s commitment to honest measurement—every future tracker must be benchmarked against the same seed to ensure meaningful comparisons. As they state, “vendors who show only successes ask for faith; a published failure matrix asks for measurement.”
From an engineering perspective, v2 is optimized for real-time operation, averaging around 1.2ms per sensor tick at the maximum density of 400 objects, with a worst-case of approximately 5ms against a 10ms budget. This performance enables browser-based real-time testing and validation. Anyone can reproduce every row of this benchmark simply by visiting the live demo and clicking “Run benchmark”—no signup or NDA required.

The entire testing harness is built on a fully synthetic environment—no real-world data, vehicles, or persons. Every pixel is generated to ensure perfect ground truth, making the benchmark a rigorous, honest measurement of the tracker’s capabilities. The use of an AI executor to build v2 against a written acceptance contract, independently reviewed, underscores the engineering discipline behind this project.
As a community committed to software quality and validation, it’s vital to recognize the value of such detailed, reproducible benchmarks. They demonstrate that even state-of-the-art systems still face thousands of identity errors under stress, highlighting areas for future improvement. We encourage readers to explore the public benchmark and try reproduce it live themselves to see the engineering rigor firsthand.
real-time object tracker software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
synthetic scene benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
video tracking performance monitor
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.