WilWell Technologies CivicAttest™

How to test whether an ALPR audit tool actually works.

Camera vendors and independent practices alike now offer tools that claim to detect misuse of license plate reader systems, which is a genuine improvement over having nothing at all. This page sets out how anyone can check whether such a tool does what it claims, and we hold our own engine to every test on it.

A detection tool is a claim about the future, since it asserts that if misuse occurs, the tool will surface it. Claims like that are only worth what the evidence behind them is worth, and the evidence has to exist before anyone relies on the tool, rather than arriving after a case goes wrong. What follows is not a criticism of any product, because we would rather every tool in this category get better, and a shared standard is how that usually happens.

Why this question is worth asking out loud.

Auditing a plate reader system is unusual work, because the volume defeats the ordinary method. A department can generate hundreds of thousands of searches in a year, which makes reading every one impossible, so oversight has historically meant a manual spot check of a small random sample or a reaction to a complaint that someone already filed. Automated detection is a real answer to that problem, and the reason to test it carefully is precisely because so much now rests on it.

The difficulty is that a detection tool almost always looks like it is working. It produces some findings, and the findings look plausible, and nobody can tell from the output alone whether the tool caught most of what was there or a small fraction of it. The only way to know is to run the tool against activity where the right answer is already known.

These are the nine tests.

Test one

Planted patterns with a sealed answer key.

Someone other than the tool's author writes synthetic search logs containing specific misuse behaviors, records exactly which records constitute each behavior, and seals that key. The tool then runs and its results are locked before the key is opened. Without a key written in advance, every result can be rationalized after the fact, and a tool that finds four things out of eight looks identical to a tool that found everything there was.

Test two

Trap scenarios that must not fire.

The same logs carry legitimate police work engineered to look statistically suspicious, meaning several officers converging on one plate during a stolen vehicle call, or a cluster of overnight searches during a missing person response. A tool that flags these is worse than useless, because it teaches a department that the alerts are noise, and the alert everyone ignores is the one that mattered.

Test three

A false accusation count, stated as a number.

The trap rows make this measurable rather than rhetorical, so a tool should be able to say how many times it pointed at legitimate work across its whole test program. The honest form of this figure is a count and a denominator rather than an adjective, and a tool whose publisher cannot produce it has not measured the thing that matters most to the officer on the receiving end.

Test four

Every miss disclosed, not only the catches.

A test program that reports only what a tool found is marketing rather than evidence. Publishing the misses tells a reader what the tool does not see, which is the single most useful thing an oversight body can know before relying on it, and it is also the part that no vendor enjoys writing.

Test five

Determinism, so the same file always produces the same findings.

An audit that returns different results on different runs cannot be checked by anyone, and it cannot be defended in a hearing where somebody asks why the number changed. Deterministic arithmetic over the records makes a finding reproducible by a third party, which is what turns a report into evidence.

Test six

Every finding traceable to the exact records behind it.

A flag that says an account behaved unusually, without naming which searches, gives a reviewer nothing to work with and gives the accused nothing to answer. Each finding should list its underlying records so a supervisor can read them, and so the person flagged can show the duty roster or the case file that explains them.

Test seven

An innocent explanation printed beside every finding.

Nearly every pattern that indicates misuse also occurs for ordinary reasons, since overtime, a call-out, a long investigation, and a repeat offender all produce activity that looks irregular in a log. A tool built for accountability rather than accusation states the innocent explanation alongside the finding, because the purpose is to raise a question a human can answer rather than to deliver a verdict nobody reviewed.

Test eight

Tested at the size of a real export, and honest about the ceiling.

A tool validated only on small samples can fail quietly on a full year from a busy department, either by breaking outright or, far worse, by silently disabling part of its own logic at scale. Testing should happen at the record counts real agencies actually produce, and whatever limit exists should be published rather than discovered by a client.

Test nine

A reporting path that does not end inside the audited agency.

This one is structural rather than technical, and it is the test a vendor built tool cannot pass on its own. When detection reports only to an administrator inside the department under review, the department decides what happens next, and the public is asked to accept that the process worked. That may well be true, and the point is that nobody outside can confirm it. A tool intended to reassure a community needs a path by which findings reach someone the community can see, which usually means a published version and a way to verify the document has not been altered.

Here is how our own engine measures against this.

We publish our record because a standard nobody applies to themselves is not worth reading. Our detection engine has been measured across 1.6 million synthetic records, comprising an original forty log program and three separate tests of roughly 500,000 records each that were run blind, with results locked to disk before any sealed answer key was opened. Across that program, 164 planted patterns were scored and 127 were detected, the accusatory rules touched zero trap rows representing legitimate stolen vehicle and missing person responses, and every finding attributed to a planted behavior in the large file tests contained only planted records.

The misses are published in the same place as the results, and there are three behavior classes our engine did not catch in the large file tests, which we describe along with the mechanism that caused each one. The complete record, including the failures, is in our accuracy report.

Publishing the tests is not the same as publishing the trigger points. A tool whose exact numeric thresholds are public tells anyone who wants to avoid it how far to stay under them, so our thresholds are disclosed under confidentiality to a client's attorney or a reviewer of the client's choosing rather than posted here. Every test above can be run and verified without knowing a single threshold value, which is the point of writing them this way.

What this standard does not settle.

Passing these tests says a tool detects what it was built to detect, and it says nothing about whether the underlying surveillance program should exist, which is a decision for a community and its elected officials rather than for an auditor. It also does not address retention schedules, destruction records, or whether data sharing matched a signed agreement, since those questions live outside a search log and need a broader data practices review. We would rather state those limits plainly than let a clean audit be read as a broader endorsement than it is.

This is offered for anyone to use.

Any vendor, agency, oversight board, or auditor is welcome to apply these tests, including to us, and we will answer questions about our methodology from anyone evaluating it seriously. If you build a tool in this category and you can beat our numbers, publish them the way we have published ours, because a community deciding whether to trust a plate reader program deserves more than competing assurances.

You can read the full accuracy record, see what a complete export makes possible, or check any report we issue on the verification page.

CivicAttest™ · an independent audit by WilWell Technologies civicattest.com