Research · accuracy benchmark

AI Product Data Enrichment Accuracy Benchmark

Accuracy is defined before results are calculated. A value is accurate only when it matches the verified ground truth under the documented comparison rule, including any permitted normalization or numeric tolerance.

Accuracy is not completeness. A record can be highly complete and still contain wrong values, or highly accurate while leaving many fields unresolved.

Current technical benchmark: 61 real products processed and 242 useful technical attributes produced in the current evaluation run. Public reporting focuses on observable output quality, supported product facts and conflict handling; proprietary model and decision logic remain internal.
Accuracy Contract SurfaceDEFINED BEFORE MEASUREMENT
COMPLETENESS AND ACCURACY ARE REPORTED SEPARATELY
VERIFIED GROUND TRUTH

Reference value

Built independently from approved evidence and expert review.

AI OUTPUT

Candidate value

Compared under the published rule. No result is accepted because it looks plausible.

Accuracy decisionMatch to ground truth under stated normalization and tolerance rules.
Permitted normalizationUnit conversion and format normalization may be allowed when documented.
Numeric toleranceDeviation inside the stated tolerance can be treated as correct.
Unresolved referenceReported as unresolved rather than forced into the accuracy score.
Dataset

Build the benchmark around the cases most likely to break enrichment.

A credible benchmark should not be a collection of easy products. It needs identity variation, different attribute families, multilingual naming, conflicting evidence and records where the correct action is to leave a field unresolved.

Benchmark Dataset ContractLOCK BEFORE SCORING
IDENTITYATTRIBUTE LOGICSOURCE VARIANCE
STRONG IDsCATEGORY SHIFTCONFLICTSLANGUAGE / FORMAT
Lock the sampleFreeze product IDs, truth fields, category schemas and inclusion rules before evaluating the system.
Stratify by difficultyInclude strong identifiers, weak supplier records, multiple attribute families, multilingual labels and contradictory sources.
Score unresolved correctlyA field intentionally held because evidence conflicts is not the same failure as publishing a wrong value.
Publish denominatorsFor every metric publish the number of products, truth fields, predicted fields and excluded records.
Public seed referencesAC820825 fan impeller · 00169811A17 tractor-compatible hitch · 0000656410KR combine knife · KK053090 reversible wear point. These four verified public records illustrate benchmark strata; they are not presented as a statistically sufficient accuracy benchmark.
Ground truth

Build the reference independently of the AI result.

An accuracy benchmark is only meaningful when the reference set is fixed before scoring and reviewers do not derive truth from the model output being evaluated.

Ground-Truth Construction ProtocolINDEPENDENT REFERENCE
01 · Freeze the product setRecord product identity, source context and target schema before evaluation.
02 · Build truth from evidenceUse manufacturer documentation, reliable source records and expert review where required.
03 · Label each target fieldReference value, unit, concept and product relationship must be explicit.
04 · Lock the truth setHash/version the reference so the model result cannot change the denominator.
05 · Score blind to marketing claimsCompare output against truth with deterministic error classes.
06 · Publish the denominatorEvery percentage must state products, fields and exclusions.
AI output remains outside the reference-building process. Unresolved reference fields stay unresolved rather than being invented to complete the benchmark.
Accuracy calculation

Score at field level and publish the denominator.

Do not report “95% accuracy” without defining what was scored. Separate correct values from missing, wrong-value, wrong-unit, wrong-concept and wrong-product errors.

Attribute Accuracy LedgerFORMULA + ERROR CLASSES
CORRECTValue matches truth after permitted normalization.count_correct
MISSINGRequired truth field absent from output.count_missing
WRONG VALUECorrect concept, incorrect value.count_wrong_value
WRONG UNITQuantity/unit conversion or unit identity wrong.count_wrong_unit
WRONG CONCEPTSource fact bound to the wrong attribute.count_wrong_concept
WRONG PRODUCTEvidence or value belongs to another product/variant.count_wrong_product
FIELD ACCURACYcorrect truth fields ÷ all scored truth fieldsPublish: N products · N truth fields · date · exclusions
Attribute types

Report accuracy by field type instead of hiding risk inside one average.

Identifier, numeric, unit, categorical, text and compatibility errors have different business impact. The benchmark should publish a separate denominator and error mix for each class.

Attribute Risk MatrixPUBLISH N + ERROR MIX FOR EACH CLASS
HIGH RISK
IdentifiersSKU / MPN / GTIN / OEM context. One character can resolve the wrong product.REPORT: correct / scored
HIGH RISK
Numeric dimensionsFit-critical quantities and clearances.REPORT: value + unit errors
HIGH RISK
CompatibilityProduct-to-machine relationships require identity evidence.REPORT: wrong-product separately
MEDIUM RISK
Categorical fieldsMaterial, type, class and controlled vocabularies.REPORT: wrong-concept separately
MEDIUM RISK
Free-text factsClaims must remain traceable to accepted source evidence.REPORT: unsupported claims
NORMALIZATION
Units / terminologyEquivalent representations should resolve without changing meaning.REPORT: normalization errors
Common errors

Publish failure modes with counts, not only one percentage.

A benchmark becomes operational when teams can see how the system fails and which errors create downstream risk.

Failure Taxonomy MapCOUNT EACH CLASS
Error eventone scored field receives one primary failure class
Wrong product / variantEvidence belongs to another entity.highest identity risk
Wrong concept bindingCorrect text mapped to the wrong canonical attribute.semantic risk
Wrong valueCorrect field, incorrect extracted value.factual risk
Wrong unitQuantity or conversion is incorrect.technical risk
MissingTruth field not recovered.coverage risk
Unsupported extraOutput adds a fact outside the truth/evidence contract.hallucination risk
Publish frequency as count(error class) ÷ all scored fields, with N and date. Do not fill the chart from intuition.
Confidence

Do not call an internal confidence label a calibrated probability.

ENRIVAQ uses confidence and routing signals operationally. A public calibration claim requires a separate benchmark comparing confidence bands against observed correctness on a locked truth set.

Confidence Calibration ContractWHAT MUST BE MEASURED
BAND Acount scoredobserved correct / total
BAND Bcount scoredobserved correct / total
BAND Ccount scoredobserved correct / total
BAND Dcount scoredobserved correct / total
Calibration testIf a band is presented as “high confidence”, publish its observed error rate and denominator. Until then, confidence remains an operational routing signal — not probability.
Review routing

Public policy should explain the evidence state, not expose arbitrary internal thresholds.

There is no single universal AUTO / REVIEW / REJECT number for every field. Routing depends on identity, provenance, canonical mapping, conflict state and field risk.

Evidence-to-Decision TerrainSTRUCTURAL ROUTING
CONTINUE
Supported evidenceProduct identity, provenance and binding are sufficient; no unresolved conflict.eligible for normal path
REVIEW / HOLD
Ambiguous evidenceWeak identity, incomplete provenance, competing mappings or conflicting values.do not silently save
REJECT
Invalid evidenceWrong product, listing/noise, prohibited content or hard conflict.exclude from accepted record
Benchmark review share only after a dated run: reviewed fields ÷ all evaluated fields, using the same routing policy.
Limitations

State what was tested — and what was not.

Accuracy varies with source quality, category, schema complexity, language, identifier availability and the tested system version. One dataset is not a promise for every future catalog.

Benchmark Scope BoundaryNO INDUSTRY-AVERAGE CLAIM
Tested scopeResults belong to the documented benchmark conditions.
SOURCE
Source qualityWeak or outdated sources cap achievable accuracy.
CATEGORY
Product categoryVariant-heavy families are harder to identify correctly.
SCHEMA
Schema complexityStricter schemas expose more failure points.
LANGUAGE
LanguageSource language affects extraction and review.
IDENTIFIERS
Identifier availabilityWithout identifiers, identification may limit everything else.
VERSION
System versionResults are tied to the tested workflow/model version.
NOT AN INDUSTRY AVERAGE · NOT A FORWARD GUARANTEE
Methodology

Publish enough detail to understand the result — without publishing the product recipe.

Public benchmark reporting should state the evaluation scope, what counts as a useful technical attribute, the main quality dimensions and representative error classes. Internal scoring rules, thresholds, source-selection logic and model-routing details remain proprietary.

Benchmark Control SurfaceINTERNAL CONTROLS · PUBLIC SUMMARY
BENCHMARK CONTROL 01

Evaluation scope

The benchmark uses a fixed technical-product evaluation set and reviewed reference data.

statusCONTROLLED
purposeKeeps benchmark runs comparable over time.
INTERNAL QUALITY CONTROL
Public reporting gives buyers the benchmark scope, observable results and representative examples without exposing proprietary source-selection, routing or model logic.
Keep reading

Validation, confidence and related benchmark pages.

Benchmark Evidence RouteRELATED METHODS
01
ValidationProduct data validation
02
ConfidenceAI confidence scores
03
AI validationValidate AI-generated product data
04
ComparisonAI vs manual benchmark
05
Current run61 products · 242 useful technical attributes

Talk to ENRIVAQ

Request a catalog assessment

Tell us enough to make the next step useful for your catalog.