Why the document model matters
WooCommerce product data is designed for transactions and administration. Vespa is designed for retrieval, filtering, ranking, and serving results at low latency. A successful integration starts by translating the product catalog into documents that match how customers search and how the storefront filters products.
A useful product document should contain:
- A stable product identifier from WooCommerce
- Searchable text such as the product name, short description, and description
- Structured attributes such as price, stock status, brand, and categories
- Facet fields such as color, size, material, and product type
- Ranking signals such as sales count, review score, and popularity
- URLs and display fields needed by the storefront
Do not copy the WooCommerce database schema one-for-one. Store fields that support the user experience, and normalize values before they are sent to Vespa.
Map WooCommerce entities to Vespa documents
The simplest model is one Vespa document per WooCommerce product. Use the WooCommerce product ID as the external identity:
WooCommerce product ID: 12345
Vespa document ID: product-12345
A typical document can represent a simple product, a variable product, or a grouped product. The parent product can contain fields common to all variations, while variation-specific values can be represented separately when customers need to search or purchase a particular variation.
For many catalogs, the following mapping works well:
| WooCommerce data | Vespa representation |
|---|---|
id |
product_id attribute and summary field |
name |
Indexed text field |
slug |
Summary field and attribute |
short_description |
Indexed text field |
description |
Indexed text field |
regular_price and sale_price |
Numeric attributes |
stock_status |
Attribute used for filtering |
stock_quantity |
Numeric attribute when inventory is available |
categories |
Array of strings or normalized category IDs |
| Product attributes | Arrays of normalized strings |
average_rating |
Numeric ranking attribute |
total_sales |
Numeric ranking attribute |
images |
Summary fields or an array of image URLs |
Keep the original WooCommerce ID even if the product slug changes. Slugs are useful for URLs, but they are not reliable primary keys.
Define a product schema
The following schema provides a practical starting point for product search, filtering, and result rendering:
schema product {
document product {
field product_id type string {
indexing: attribute | summary
attribute: fast-search
}
field name type string {
indexing: index | summary
}
field short_description type string {
indexing: index | summary
}
field description type string {
indexing: index | summary
}
field slug type string {
indexing: attribute | summary
}
field brand type string {
indexing: attribute | summary
attribute: fast-search
}
field categories type array {
indexing: attribute | summary
attribute: fast-search
}
field attributes type array {
indexing: attribute | summary
attribute: fast-search
}
field price type double {
indexing: attribute | summary
}
field regular_price type double {
indexing: attribute | summary
}
field on_sale type bool {
indexing: attribute | summary
}
field stock_status type string {
indexing: attribute | summary
attribute: fast-search
}
field stock_quantity type int {
indexing: attribute | summary
}
field average_rating type double {
indexing: attribute | summary
}
field review_count type int {
indexing: attribute | summary
}
field total_sales type int {
indexing: attribute | summary
}
field product_url type string {
indexing: summary
}
field image_url type string {
indexing: summary
}
}
fieldset default {
fields: name, short_description, description, brand, categories, attributes
}
rank-profile default {
first-phase {
expression: nativeRank
}
}
}
index makes text searchable. attribute makes a field available for filtering, sorting, grouping, and ranking. summary makes the value available in the document response. A field can use more than one of these indexing modes.
For high-cardinality identifiers, attribute: fast-search is useful when queries frequently test exact membership, such as a product ID or category ID. Avoid adding fast search to every field without measuring memory use.
Normalize WooCommerce values before feeding Vespa
WooCommerce data often contains HTML, inconsistent casing, formatted prices, and taxonomy values with different names. Normalize these values in the ingestion pipeline rather than relying on query-time cleanup.
For example:
{
"id": 12345,
"name": "Merino Wool Hiking Socks",
"regular_price": "24.00",
"sale_price": "19.00",
"stock_status": "instock",
"categories": ["Hiking", "Socks"],
"attributes": {
"pa_material": ["Merino wool"],
"pa_size": ["Medium", "Large"]
}
}
Can become:
{
"fields": {
"product_id": "12345",
"name": "Merino Wool Hiking Socks",
"short_description": "Breathable socks for hiking and trekking.",
"description": "Merino wool hiking socks with reinforced heel and toe.",
"slug": "merino-wool-hiking-socks",
"categories": ["hiking", "socks"],
"attributes": [
"material:merino-wool",
"size:medium",
"size:large"
],
"price": 19.0,
"regular_price": 24.0,
"on_sale": true,
"stock_status": "instock",
"stock_quantity": 87,
"product_url": "https://store.example.com/product/merino-wool-hiking-socks/"
}
}
Use a consistent facet format such as attribute:value. This prevents a value such as Large from being confused with a category or another attribute. Normalize case, whitespace, punctuation, and taxonomy slugs consistently across all products.
Strip HTML tags from descriptions while preserving useful text boundaries. Remove navigation fragments, repeated boilerplate, and shortcodes before indexing. Keep the original formatted description in another system if the storefront needs it for rendering.
Choose field types based on query behavior
Use strings for identifiers, labels, URLs, and exact categorical values. Use numeric types for values used in ranges, sorting, or calculations.
For example, price should be a double, not a formatted string such as "$19.00". This allows queries such as:
range(price, 10, 25)
and ordering such as:
order by price asc
Currency requires an explicit decision. If the catalog uses one currency, a numeric price field is usually sufficient and the currency can be handled by application configuration. If products from multiple currencies are indexed together, store a currency field and either convert prices to a common comparison currency or keep separate document sets and query rules. Do not compare raw numeric values from different currencies.
Boolean fields are useful for common states such as on_sale, featured, or backorders_allowed. Use strings for WooCommerce statuses when there are several named states, such as instock, outofstock, and onbackorder.
Model categories and product attributes for faceting
Categories and attributes have different meanings. Categories describe where a product belongs in the catalog. Attributes describe characteristics that shoppers may use to narrow results.
A product might have:
{
"categories": ["footwear", "hiking-boots"],
"attributes": [
"size:42",
"size:43",
"color:brown",
"material:leather",
"waterproof:true"
]
}
At query time, an application can apply filters such as:
attribute(attributes contains "color:brown")
and attribute(stock_status contains "instock")
and attribute(price in range [80, 180])
The exact YQL syntax should be generated and escaped by the application rather than concatenated from untrusted request parameters. For large catalogs, storing category and attribute IDs can reduce ambiguity and make taxonomy changes easier to manage:
categories: ["term-18", "term-42"]
attributes: ["color:term-7", "size:term-11"]
Resolve those IDs to display labels in the application or include a separate summary field when the response must be self-contained.
Handle variable products and variations
Variable products require a modeling decision. A parent product may have a shared name, description, images, and category, while each variation has its own SKU, price, stock state, and attributes.
Use one document per parent product when search results should show one result for the product. Aggregate variation data into fields such as:
{
"fields": {
"product_id": "12345",
"attributes": [
"color:black",
"color:blue",
"size:medium",
"size:large"
],
"price": 29.0,
"stock_status": "instock"
}
}
This is simple, but it cannot accurately represent every variation’s price and availability. Use separate variation documents when customers search for exact SKU-level inventory, when variation prices differ substantially, or when the purchase flow needs a direct variation identifier. A variation document can include parent_product_id, sku, price, stock_status, and its own normalized attributes.
If both parent and variation documents are indexed, add a document type or is_variation field and filter the result set appropriately. Otherwise, search results may contain both the parent and its child variations.
Feed documents with stable IDs
A document feed should use a deterministic document ID. For example:
id=store.product::12345
A feed request to the document API can look like this:
curl --request PUT \
--header 'Content-Type: application/json' \
--data '{
"fields": {
"product_id": "12345",
"name": "Merino Wool Hiking Socks",
"price": 19.0,
"stock_status": "instock",
"categories": ["hiking", "socks"],
"attributes": ["material:merino-wool", "size:medium"]
}
}' \
'https://vespa.example.com/document/v1/store/product/docid/12345'
The exact host and deployment endpoint depend on the Vespa environment. The document type in the URL must match the schema name, and the document ID must remain stable across updates.
Use full document updates when the feed produces a complete canonical representation. For small, frequent changes such as stock or price, partial updates can reduce payload size, but the update path and field definitions must support the operation. Inventory updates should be idempotent so that retries do not create inconsistent state.
Keep catalog synchronization event-driven
A WooCommerce integration should normally combine an initial bulk import with incremental updates. Import all published products first, then process product, price, inventory, taxonomy, and deletion events.
At minimum, handle:
- Product creation and updates
- Product deletion or transition to an unpublished status
- Variation updates
- Stock changes
- Price and sale-period changes
- Category and attribute-term changes
When a product is deleted or should no longer be searchable, issue a document delete rather than feeding an empty product. During a full reindex, use a new Vespa deployment or controlled document namespace when appropriate, validate counts and sample queries, and switch traffic only after the replacement index is ready.
A common operational mistake is updating price and stock only in WooCommerce while leaving Vespa stale. Add freshness monitoring for these fields and alert when the age of the latest successful catalog event exceeds the business requirement.
Add ranking signals deliberately
Text relevance should identify products that match the query. Business signals can then influence ordering among relevant products. Useful signals include review score, review count, sales volume, margin, freshness, and inventory state.
A simple rank profile can expose these values as rank features:
rank-profile product-ranking inherits default {
first-phase {
expression: nativeRank
+ 0.15 * log(1 + attribute(total_sales))
+ 0.10 * attribute(average_rating)
+ if(attribute(stock_status) == "instock", 0.2, -1.0)
}
}
The exact expression should be validated against real search results. Do not let popularity completely override textual relevance, and do not use a stock boost as a substitute for filtering products that cannot be purchased.
If ranking depends on a value that changes frequently, such as inventory or click-through rate, consider how often it is updated and whether the added operational complexity is justified. Rank profiles can evolve independently from the ingestion code, but field names and types must remain compatible.
Return only the fields the storefront needs
Define document summaries for search responses rather than returning every indexed field. A result summary commonly includes the product ID, name, URL, image URL, price, sale state, rating, and selected facet values.
For example, a response-oriented product result might contain:
{
"product_id": "12345",
"name": "Merino Wool Hiking Socks",
"price": 19.0,
"on_sale": true,
"stock_status": "instock",
"image_url": "https://store.example.com/images/socks.jpg",
"product_url": "https://store.example.com/product/merino-wool-hiking-socks/"
}
Keep large descriptions, raw metadata, and internal synchronization fields out of the default summary unless the result page needs them. Smaller responses reduce latency and bandwidth, especially for autocomplete and category pages.