Developers · Architecture

Product Data Enrichment Architecture

ENRIVAQ is designed as an enrichment layer between incomplete source data and the systems that ultimately manage or publish the product record. The architecture separates ingestion, research, AI processing, validation and export so each stage has a clear role in the workflow.

This is a high-level public architecture. It explains data movement and processing boundaries without exposing infrastructure secrets, credentials, internal hostnames or private network details.

STAGE OBSERVATORY9 WORKFLOW STAGES · CROSS-CUTTING CONTROLS
STAGE 01IngestionExisting records enter and are mapped to a predictable structure.
STAGE 02OrchestrationCoordinates the stages each record needs.
STAGE 03ResearchCollects evidence that may support missing fields.
STAGE 04AI processingExtraction, classification, normalization and content tasks.
STAGE 05ValidationConflicts, gaps and uncertainty are controlled here.
STAGE 06StorageJob state, product records, source context and results.
STAGE 07ExportValidated records move to the owning downstream system.
CROSS-CUTTINGSecurity boundariesWhere data enters, is processed and leaves.
CROSS-CUTTINGMonitoringJob health, failures and exceptions.
PUBLICATION RULE: high-level architecture only. No private deployment topology, credentials, internal hostnames or unverified controls.
High-level system diagram

The main processing boundaries, not the deployment

The public diagram shows where data enters, which stages act on it and where a validated record leaves. Internal deployment topology stays out of the published version.

BOUNDARY BLUEPRINT CUTAWAYPUBLIC WORKFLOW · VERIFIED CAPABILITIES
INPUT
Existing catalog recordsCustomer-provided product context.
PROCESSING
IngestionMap incoming records.
OrchestrationCoordinate required stages.
ResearchCollect candidate evidence.
AI processingRun approved model-assisted tasks.
CONTROL
Validation / reviewHold, reject, review or approve.
Structured product recordValidated working result.
OUTPUT
ExportMove result to supported destination.
Connected catalog / configured site export where enabledNo generic native PIM/ERP/ecommerce connector or file-export destination is implied.
PUBLIC FLOW: Website / API / File → Identify → Research → Extract → Normalize → Validate → Content / SEO → Connected catalog / configured site export where enabled.
Stage 01 · Ingestion

Where existing records enter the workflow

The platform accepts only input methods that are actually supported in production. Each record is mapped into a predictable structure before enrichment begins.

Exact formats, field limits and endpoint behavior belong in Import & Export and API documentation, which remain authoritative.

INGRESS GATE / SCHEMA MAPPERSUPPORTED METHODS ONLY
SOURCE RECORD

Context in

Product identifiersREQUIRED
Existing attributesCONTEXT
CategoryCONTEXT
Supported transportPRODUCTION ONLY
VALIDATE
+ MAP
INTERNAL SHAPE

Predictable structure

Malformed recordREJECT / FLAG
Mapped fieldsREADY
Identity contextBOUND
Next stageORCHESTRATE
Stage 02 · Orchestration

One record, several coordinated stages

A record may need identification, research, extraction, normalization, validation, content generation and review. These stages are not presented as one opaque AI call.

JOB CHOREOGRAPHYPUBLIC WORKFLOW · INTERNAL RETRY DETAILS NOT PUBLISHED
IdentificationResolve the real product.
ResearchFind usable evidence.
ExtractionRead structured values.
NormalizationAlign units and naming.
ValidationCheck gaps and conflicts.
ContentBuild from validated data.
ReviewEscalate uncertainty.
Public workflow: records move through identification, research, extraction, normalization and validation; supported data continues, ambiguity/conflicts are held and invalid/irrelevant evidence is rejected. Exact retry/state internals are not published.
EVIDENCE HARVESTEVIDENCE, NOT TRUTH
PRODUCT CONTEXT

Identifiers + existing attributes + category

This context is used to search for documentation that plausibly describes the same product.

CANDIDATEManufacturer / technical documentationCollected when it plausibly matches the product context.
CANDIDATESupporting product sourcesUseful excerpts and references are attached as evidence.
OUTPUTSource contextEvidence passes to extraction and validation.
Research never writes a final value. Extraction and validation decide; proprietary ranking logic stays private.
Stage 03 · Research

Research collects evidence for the next stage

Research output is not treated as final truth. It becomes evidence for extraction and validation.

The public architecture describes the role of research without exposing proprietary source-selection or ranking logic.

Stage 04 · AI processing

Model-assisted tasks remain inside defined boundaries

AI may assist identification, extraction, classification, normalization and content generation where those capabilities are actually implemented. Deterministic checks, schema mapping and export remain outside the model.

MODEL BOUNDARY SWITCHBOARDPUBLIC BOUNDARY · PRIVATE PROVIDER ROUTING NOT DISCLOSED
MODEL-ASSISTEDProduct identificationResolve a reference to the actual product.CONTEXT → CANDIDATE MATCH
MODEL-ASSISTEDExtractionRead attribute values from source material.SOURCE → FIELD VALUES
MODEL-ASSISTEDClassificationAssign to the taxonomy in use.RECORD → TAXONOMY
MODEL-ASSISTEDNormalizationAlign units, notation and naming.RAW → NORMALIZED
MODEL-ASSISTEDContent generationBuild content from validated attributes.VALIDATED DATA → CONTENT
DETERMINISTICRules stay outsideRule checks, conversions, schema mapping and export use no model.NO MODEL INVOLVED
Model/provider categories and customer-data handling are maintained on the AI & Data Policy page; private routing and provider-selection details are not duplicated here.
Stage 05 · Validation

The control layer before a record becomes final

Validation identifies missing required fields, structural problems, source conflicts, uncertainty and records that require human review.

A record that fails validation does not silently continue. It is held, flagged or routed to review, and unresolved fields stay empty.

DECISION GATEHOLD · REJECT · REVIEW · READY
Missing required fieldsCompare against the required field set for the assigned category.
HELD
Structural problemsType, unit, format and schema conformance.
REJECTED
Conflicting sourcesDifferent values supporting the same attribute.
REVIEW
Uncertain valuesWeak source support or low-confidence match.
REVIEW
Passed recordComplete, conformant and supported by source reference.
READY
Stage 06 · Storage

State and evidence stay inspectable between stages

The working architecture may store job state, product records, source context, stage outputs and operational logs. Retention, encryption and region claims are not published here until verified.

STATE LEDGER VAULTPOLICY DETAILS → DATA PRIVACY / SECURITY / DPA
STATEJob stateTrack where each record is and what failed.See Data Privacy / DPA for the applicable retention terms.
RECORDProduct recordsWorking copy being enriched and validated.See Data Privacy / DPA for the applicable retention terms.
EVIDENCESource contextReferences and excerpts supporting values.See Data Privacy / DPA for the applicable retention terms.
OUTPUTProcessing resultsStage outputs for traceability.See Data Privacy / DPA for the applicable retention terms.
OPSOperational logsDiagnose failures and audit actions.See Data Privacy / DPA for the applicable retention terms.
Stage 07 · Export

Validated records move back to the system that owns the workflow

Public architecture language stays simple: existing system → ENRIVAQ enrichment → validated record → downstream system.

EGRESS FAN-OUTDESTINATIONS = SUPPORTED METHODS ONLY
VALIDATEDProduct record
Customer catalogCONFIGURED HANDOFF
Existing stackCONFIGURED HANDOFF
Configured siteENABLED DESTINATION
File destinationWHEN CONFIGURED
External API resultDEPLOYMENT CONTRACT
ASYNCHRONOUS STATE LANESCONCEPTUAL WORKFLOW · NOT PUBLIC API ENUM
RECEIVEDInput accepted by the configured workflow.
PROCESSINGRequired enrichment stages are running.
EXCEPTIONA processing or evidence issue requires handling.
REVIEW / HOLDAmbiguous or conflicting evidence is kept out of accepted output.
RESULT READYAccepted output is available for the configured handoff.
FAILEDThe configured processing path did not complete.
Queues and jobs

Catalog work is a job, not a prompt

Catalog enrichment is a workflow across many products, not a sequence of individual chat prompts. Where jobs or queues are used, processing is asynchronous and job state can be tracked.

Exact queue names, retry counts, concurrency and external state names are implementation details; any public API state contract is configured.

Security boundaries

Where data enters, where it is processed, where it leaves

Boundaries are shown conceptually. Detailed controls belong on the Security page and describe only measures that are actually implemented.

TRUST BOUNDARY CROSS-SECTIONCONTROL DETAILS LIVE ON SECURITY PAGE
B1Where data entersAuthenticated ingestion through supported files, connectors or API endpoints.CUSTOMER → PLATFORM
B2Where AI processing may occurModel calls happen inside a defined boundary with a defined payload.PLATFORM → MODEL
B3Where validation happensInside the platform before a record is considered publishable.PLATFORM INTERNAL
B4Where results leaveExport to systems and formats configured by the customer.PLATFORM → DOWNSTREAM
Scalability

Scale comes from controlled work distribution, not one oversized request

The public architecture can explain batch-first workflow, stage isolation, retry safety and controlled concurrency. Real throughput figures are published only when verified.

WORKCELL MATRIXNO PUBLIC PRODUCTION FIGURES WITHOUT RUN LEDGER
01Batch-first workflowWork is defined per catalog segment, not one interactive request.
02Stage isolationA slow or failing stage does not need to block unrelated work.
03Idempotent retriesA repeated attempt should not create a duplicate record.
04Controlled concurrencyThroughput is limited deliberately so quality controls keep up.
PROCESSED SKUNo public production figure without a run ledger
CONCURRENT JOBSNo public production figure without a run ledger
APPROVED METRICNo public production figure without a run ledger
Monitoring

Failures are visible to the team running the platform

Monitoring makes processing failures, job health and exceptions visible to the operating team. Operational tooling and alerting details remain internal.

Alert thresholds, on-call rotation and internal dashboards are not published here.

OBSERVABILITY RADARPUBLIC OPERATIONAL VIEW
Job healthQueue depth, stage duration and stalled records.INTERNAL OPERATIONAL DETAIL
Processing failuresFailed stages grouped by cause.INTERNAL OPERATIONAL DETAIL
ExceptionsUnexpected conditions raised by a stage.INTERNAL OPERATIONAL DETAIL
Integration healthConnector and API errors at the ingestion/export edges.INTERNAL OPERATIONAL DETAIL

Talk to ENRIVAQ

Request a catalog assessment

Tell us enough to make the next step useful for your catalog.