Table of contents :

Designing a Vespa Schema for WooCommerce Products

100-days-of-vespa-ai-woocommerce-004

Table of contents :

Start with the WooCommerce search model

A Vespa schema should reflect how customers and downstream services search WooCommerce data, not simply mirror the WordPress database tables. For most stores, the searchable unit should be a product or a purchasable variation. Each document should contain the text used for relevance, the structured values used for filtering, and the identifiers required to link results back to WooCommerce.

A practical document model commonly includes:

  • A stable product or variation identifier
  • Parent product and variation identifiers
  • Title, description, short description, and searchable attributes
  • SKU and optional brand or product-type fields
  • Price and stock status
  • Category and taxonomy identifiers
  • Facet values such as color, size, material, and brand
  • A product URL or canonical permalink
  • Optional image, rating, and popularity fields for result presentation or ranking

Keep the Vespa document independent of WooCommerce’s storage schema. The ingestion process should transform posts, post metadata, taxonomy relationships, and variation data into a search-oriented document.

Choosing the document boundary

There are two common approaches for products with variations.

One document per parent product

This approach is suitable when customers search for a product and select size or color after opening the product page. The parent document can contain aggregated variation information, such as available sizes, colors, minimum price, and maximum price.

One document per sellable variation

This approach is better when variation-level inventory, price, or availability must affect search results. Each variation becomes a Vespa document and includes a parent_id that points to the WooCommerce parent product. Search results can then be grouped or de-duplicated by parent_id in the application layer.

For agencies building catalog search, the second approach is usually more reliable when products have different prices, stock states, or searchable attributes per variation. It also prevents an out-of-stock variation from making an otherwise unavailable option appear purchasable.

A baseline Vespa schema

The following schema models one document per searchable WooCommerce product or variation. It supports full-text search, numeric price filtering, taxonomy filtering, and result summaries.

schema product {
  document product {
    field product_id type string {
      indexing: attribute | summary
      attribute: fast-search
    }

    field parent_id type string {
      indexing: attribute | summary
      attribute: fast-search
    }

    field sku type string {
      indexing: index | attribute | summary
      attribute: fast-search
    }

    field title type string {
      indexing: index | summary
    }

    field short_description type string {
      indexing: index | summary
    }

    field description type string {
      indexing: index | summary
    }

    field brand type string {
      indexing: index | attribute | summary
      attribute: fast-search
    }

    field product_type type string {
      indexing: attribute | summary
      attribute: fast-search
    }

    field category_ids type array {
      indexing: attribute | summary
      attribute: fast-search
    }

    field attribute_values type array {
      indexing: index | attribute | summary
      attribute: fast-search
    }

    field price type double {
      indexing: attribute | summary
    }

    field regular_price type double {
      indexing: attribute | summary
    }

    field in_stock type bool {
      indexing: attribute | summary
    }

    field stock_quantity type int {
      indexing: attribute | summary
    }

    field permalink type string {
      indexing: summary
    }

    field image_url type string {
      indexing: summary
    }

    field rating type double {
      indexing: attribute | summary
    }

    field popularity type double {
      indexing: attribute | summary
    }
  }

  fieldset default {
    fields: title, short_description, description, sku, brand, attribute_values
  }

  rank-profile lexical {
    first-phase {
      expression: bm25(title) + bm25(short_description) + bm25(description) + bm25(sku)
    }
  }

  rank-profile commercial inherits lexical {
    first-phase {
      expression: bm25(title) + bm25(short_description) + bm25(description) + bm25(sku) + 0.05 * popularity + 0.02 * rating
    }
  }
}

The schema deliberately uses different indexing modes for different field types:

  • index creates an inverted index for text matching.
  • attribute makes values available for filtering, sorting, grouping, and ranking.
  • summary makes fields available in document results.
  • attribute: fast-search improves exact-match and set-membership filtering for string attributes and arrays.

Do not add index to every field by default. A price, stock flag, or internal identifier does not need full-text indexing. Conversely, a title should normally be an indexed text field rather than only an attribute.

Mapping WooCommerce data into documents

A transformed product document might look like this when sent to Vespa:

{
  "put": "id:catalog:product::variation-1842",
  "fields": {
    "product_id": "variation-1842",
    "parent_id": "product-731",
    "sku": "RUN-SHOE-BLK-42",
    "title": "Trail Running Shoe - Black - Size 42",
    "short_description": "Lightweight trail shoe with a reinforced toe",
    "description": "Designed for mixed terrain with a durable outsole and breathable upper.",
    "brand": "North Ridge",
    "product_type": "variable",
    "category_ids": ["running-shoes", "footwear"],
    "attribute_values": ["color:black", "size:42", "terrain:trail"],
    "price": 89.95,
    "regular_price": 109.95,
    "in_stock": true,
    "stock_quantity": 14,
    "permalink": "https://shop.example.com/product/trail-running-shoe/",
    "image_url": "https://shop.example.com/wp-content/uploads/trail-running-shoe.jpg",
    "rating": 4.7,
    "popularity": 812.0
  }
}

Use a stable identifier that does not change when the product title or permalink changes. WooCommerce post IDs and variation IDs are generally suitable. The Vespa document ID can include a namespace and type, while the product_id field remains available to application code.

Normalize values before feeding them to Vespa. For example, store prices as numeric values in the store’s canonical currency rather than formatted strings such as £89.95. Store category and attribute identifiers in a stable form, and preserve display labels separately if the search interface needs them.

Modeling taxonomy and attributes

WooCommerce taxonomies can be represented in several ways. For filtering, identifiers are usually more reliable than labels because labels can be renamed or localized.

{
  "category_ids": ["men", "men-running", "footwear"],
  "attribute_values": [
    "pa_color:black",
    "pa_size:42",
    "pa_material:mesh"
  ]
}

An array field supports filters such as:

yql=where(in(category_ids, "men-running") and in(attribute_values, "pa_color:black") and in_stock)

If the user interface needs faceting by attribute name, separate fields can be easier to aggregate and present. For example, use color, size, and material as individual string attributes when those facets are central to the store’s navigation. Use a general attribute_values array for less predictable or long-tail WooCommerce attributes.

Do not combine a display label and an internal value into a field that must be filtered unless the format is controlled. A value such as Black / Noir can be displayed to customers, while a stable value such as black remains suitable for filtering.

Price, stock, and sale logic

Price and availability should be stored as typed fields. This allows Vespa to perform numeric comparisons instead of lexical comparisons.

yql=where(price <= 100 and in_stock)

Querying the schema

A basic text query using the lexical rank profile can be sent to Vespa with a request similar to:

/search/?yql=select * from sources * where userInput(@query) and in_stock&query=waterproof%20hiking%20boots&ranking=lexical

A filtered catalog query can combine text, category, and price constraints:

/search/?yql=select * from sources * where userInput(@query) and in(category_ids, "footwear") and price <= 150 and in_stock&query=trail%20shoe&ranking=commercial

The application should construct YQL using validated values rather than concatenating unchecked request parameters. Category IDs, attribute values, sort options, and numeric ranges should be validated against the catalog or an allowlist before creating the request.

Handling duplicate parent products

When variations are indexed as separate documents, one WooCommerce product can produce several search hits. There are three common strategies:

  1. Group results by parent_id and display one parent product with the best matching variation.
  2. Keep variation-level results when size, color, or stock differences are important to the customer.
  3. Use a separate parent-product document for presentation and variation documents for availability and filtering.

The correct choice depends on the storefront experience. For a quick-search dropdown, parent-level results are often easier to understand. For a technical catalog where SKU and exact dimensions matter, variation-level results may be preferable.

If parent-level display is required, store enough information to select the representative variation, such as its price, image, and availability. Avoid relying on the first document returned by Vespa without defining how the representative variation is chosen.

Feed updates and deletion handling

The indexing pipeline should process WooCommerce create, update, status, taxonomy, price, and stock changes. A product update is not limited to the post title: a change to a variation, category assignment, sale price, or inventory state can affect search results.

Use idempotent document operations. Replaying the same update should produce the same document state. When a product or variation is permanently removed, send a Vespa remove operation for the corresponding document ID. For products transitioned to draft, private, or trash status, either remove them or set a field that excludes them from customer-facing queries.

A reliable synchronization process normally includes:

  • An initial full catalog feed
  • Incremental updates from WooCommerce hooks or an integration queue
  • Periodic reconciliation against the authoritative WooCommerce catalog
  • Monitoring for feed failures and stale documents
  • A deletion or visibility policy for discontinued products

Store source timestamps or a catalog revision in the document if the pipeline needs to detect stale updates. The source system should remain authoritative for price and inventory; Vespa should not be treated as the system of record for those values.

Schema design checks for agencies

Before deploying the schema, verify the following against real catalog data:

  • Searchable text is indexed in the fields used by the rank profile.
  • Filterable values are typed consistently across all documents.
  • Prices use numeric fields and one clearly defined currency.
  • Variations have stable parent and variation identifiers.
  • Draft, private, deleted, and out-of-stock products follow an explicit visibility policy.
  • Category and attribute values remain stable when display names change.
  • Result summaries contain every field required by the storefront response.
  • Reindexing produces the same document state as incremental updates.
  • Queries are tested with products that have missing descriptions, no SKU, multiple categories, and several variations.

For a WooCommerce implementation, schema design is inseparable from feed design. The Vespa schema defines what can be searched and filtered efficiently, while the transformation layer determines whether the indexed values accurately represent the catalog customers can actually purchase.

Trending posts
You might also like