AI Product Data Enrichment Accuracy Benchmark
Accuracy is defined before results are calculated. A value is accurate only when it matches the verified ground truth under the documented comparison rule, including any permitted normalization or numeric tolerance.
Accuracy is not completeness. A record can be highly complete and still contain wrong values, or highly accurate while leaving many fields unresolved.
Reference value
Built independently from approved evidence and expert review.
Candidate value
Compared under the published rule. No result is accepted because it looks plausible.
Build the benchmark around the cases most likely to break enrichment.
A credible benchmark should not be a collection of easy products. It needs identity variation, different attribute families, multilingual naming, conflicting evidence and records where the correct action is to leave a field unresolved.
Build the reference independently of the AI result.
An accuracy benchmark is only meaningful when the reference set is fixed before scoring and reviewers do not derive truth from the model output being evaluated.
Score at field level and publish the denominator.
Do not report “95% accuracy” without defining what was scored. Separate correct values from missing, wrong-value, wrong-unit, wrong-concept and wrong-product errors.
count_correctcount_missingcount_wrong_valuecount_wrong_unitcount_wrong_conceptcount_wrong_productReport accuracy by field type instead of hiding risk inside one average.
Identifier, numeric, unit, categorical, text and compatibility errors have different business impact. The benchmark should publish a separate denominator and error mix for each class.
REPORT: correct / scoredREPORT: value + unit errorsREPORT: wrong-product separatelyREPORT: wrong-concept separatelyREPORT: unsupported claimsREPORT: normalization errorsPublish failure modes with counts, not only one percentage.
A benchmark becomes operational when teams can see how the system fails and which errors create downstream risk.
Do not call an internal confidence label a calibrated probability.
ENRIVAQ uses confidence and routing signals operationally. A public calibration claim requires a separate benchmark comparing confidence bands against observed correctness on a locked truth set.
Public policy should explain the evidence state, not expose arbitrary internal thresholds.
There is no single universal AUTO / REVIEW / REJECT number for every field. Routing depends on identity, provenance, canonical mapping, conflict state and field risk.
eligible for normal pathdo not silently saveexclude from accepted recordState what was tested — and what was not.
Accuracy varies with source quality, category, schema complexity, language, identifier availability and the tested system version. One dataset is not a promise for every future catalog.
Publish enough detail to understand the result — without publishing the product recipe.
Public benchmark reporting should state the evaluation scope, what counts as a useful technical attribute, the main quality dimensions and representative error classes. Internal scoring rules, thresholds, source-selection logic and model-routing details remain proprietary.
Evaluation scope
The benchmark uses a fixed technical-product evaluation set and reviewed reference data.
statusCONTROLLEDpurposeKeeps benchmark runs comparable over time.