Table of contents :

Choosing the Right WooCommerce Fields for Vespa

100-days-of-vespa-ai-woocommerce-005

Table of contents :

Start with a search contract

A WooCommerce product object contains far more information than most search experiences need. Before creating a Vespa schema, define which fields support each user-facing capability:

  • Search: product names, descriptions, SKUs, brands, and relevant attributes.
  • Filtering: categories, stock status, price, rating, brand, size, colour, and other faceted values.
  • Ranking: popularity, sales volume, rating, recency, margin, and inventory signals.
  • Display: the product title, URL, image, prices, availability, and selected merchandising data.

This separation prevents a common implementation problem: treating every WooCommerce property as both searchable text and a filterable field. In Vespa, field type and indexing mode affect memory usage, query behavior, ranking, and update cost.

Separate searchable text from exact-match data

Fields intended for natural-language queries should generally be indexed as text. For example, a cleaned product name and description can support queries such as waterproof hiking boots. Exact values such as a SKU, brand identifier, or stock status should normally be stored as attributes so that filtering does not depend on text matching.

field product_name type string {
    indexing: index | summary
}

field description type string {
    indexing: index | summary
}

field sku type string {
    indexing: summary | attribute
    attribute: fast-search
}

field stock_status type string {
    indexing: summary | attribute
    attribute: fast-search
}

Use summary when a value must be returned in a result. Use index for full-text search. Use attribute for efficient filtering, sorting, or ranking. A field can use more than one indexing mode when it serves multiple purposes.

Model WooCommerce identifiers deliberately

Keep the WooCommerce product ID as a stable numeric identifier, and use it as the Vespa document ID or as a separate field used during document updates. Do not use the product name or slug as the primary identity: names can change, and slugs can be regenerated.

field woo_product_id type int {
    indexing: summary | attribute
    attribute: fast-search
}

field slug type string {
    indexing: summary | attribute
    attribute: fast-search
}

field product_type type string {
    indexing: summary | attribute
    attribute: fast-search
}

For variable products, decide whether the searchable unit is the parent product, each variation, or both. Parent documents are usually appropriate when shoppers select options on the product page. Variation documents are useful when individual combinations need their own inventory, price, SKU, or search visibility. If both are indexed, add an explicit relationship such as parent_product_id and ensure that result presentation does not show duplicate products.

Use the right numeric representation for prices

WooCommerce REST responses commonly expose prices as decimal strings. Do not store those values as arbitrary text if users must filter or sort by price. Convert them during ingestion to a numeric field.

field price_minor type int {
    indexing: summary | attribute
}

field regular_price_minor type int {
    indexing: summary | attribute
}

field sale_price_minor type int {
    indexing: summary | attribute
}

Storing minor currency units, such as cents, avoids floating-point comparison issues. The conversion must use the store currency and decimal precision supplied by WooCommerce. For a multi-currency store, do not compare values from different currencies in one field. Use a currency-specific price field or normalize prices to a clearly defined comparison currency, while retaining the display currency and formatted value for the result.

Only create a sale-price field when the value is valid for the current product and sale period. A missing sale price should not be represented as zero, because zero can incorrectly make a product appear first in ascending price results.

Choose facet fields from WooCommerce taxonomies and attributes

WooCommerce categories and tags are naturally modeled as arrays. Product attributes require more care because their values may be global taxonomies, custom attributes, or variation-specific options.

field category_ids type array<int> {
    indexing: summary | attribute
    attribute: fast-search
}

field category_names type array<string> {
    indexing: summary | attribute
}

field brand type string {
    indexing: summary | attribute
    attribute: fast-search
}

field colours type array<string> {
    indexing: summary | attribute
    attribute: fast-search
}

Use stable term IDs for internal filtering when taxonomy terms can be renamed. Keep human-readable names for facet labels and result rendering. If a product has multiple values for an attribute, use an array rather than joining values into a comma-separated string. Arrays preserve the individual values and make filter expressions unambiguous.

Normalize facet values during ingestion. For example, map Blue, blue, and BLUE to one canonical value, but retain a display label if the storefront requires specific capitalization. Also decide how to handle values such as 10, 10 oz, and 10-ounce; they should not become separate facets unless that distinction is intentional.

Keep descriptions useful for search

WooCommerce descriptions often contain HTML, shortcodes, embedded links, and formatting markup. Index a cleaned text representation rather than raw editor output. Remove navigation fragments, tracking attributes, hidden content, and duplicated short descriptions where possible.

field short_description type string {
    indexing: index | summary
}

field description_text type string {
    indexing: index | summary
}

Keep the original HTML only if the application needs it for rendering, and consider storing it in a summary-only field that is not indexed. Large descriptions increase document size and can reduce the value of returning full content in every search response. A common result document contains the title, URL, image, price, availability, and a short text excerpt; the full description can be loaded from WooCommerce or a separate content source when needed.

Represent images for result rendering

Most search result cards need one primary image, while product pages may need a gallery. Store the primary image as simple summary fields and model additional images as an array or structured collection only when the result application needs them.

field image_url type string {
    indexing: summary
}

field image_alt type string {
    indexing: summary
}

field gallery_urls type array<string> {
    indexing: summary
}

Do not index image URLs as searchable text. If image ordering matters, preserve the WooCommerce image position during ingestion rather than relying on an unordered representation.

Store ranking signals independently

Search text and ranking signals should not be mixed into one field. Keep sales, ratings, inventory, and dates in fields with types that support the ranking expressions you intend to use.

field average_rating type double {
    indexing: summary | attribute
}

field rating_count type int {
    indexing: summary | attribute
}

field total_sales type int {
    indexing: summary | attribute
}

field date_modified type long {
    indexing: summary | attribute
}

field inventory_quantity type int {
    indexing: summary | attribute
}

Use a timestamp such as Unix epoch seconds for date-based ranking and freshness calculations. Treat missing inventory carefully: an unavailable quantity should not automatically become a large negative or positive value unless the ranking expression expects that behavior.

For ratings, avoid ranking solely by average rating. A product with one five-star review should not necessarily outrank a product with hundreds of consistently high ratings. A ranking expression can combine average rating with rating count, sales volume, text relevance, and availability. The exact formula belongs in the ranking profile, while the source values remain separate Vespa fields.

Handle stock and visibility as business rules

Index fields that allow the query layer to enforce storefront rules without retrieving every product first.

field catalog_visibility type string {
    indexing: summary | attribute
    attribute: fast-search
}

field stock_status type string {
    indexing: summary | attribute
    attribute: fast-search
}

field purchasable type bool {
    indexing: summary | attribute
}

The ingestion pipeline should derive these values consistently from WooCommerce status, catalog visibility, stock settings, backorders, and any agency-specific merchandising rules. For example, a product can be visible in the catalog but not purchasable, or visible only when searched directly. Encode those distinctions explicitly instead of relying on an ambiguous status string.

Use query filters for rules that must always apply, such as excluding private or deleted products. Use ranking signals for preferences, such as promoting in-stock products above backordered products. This keeps mandatory eligibility separate from optional ordering.

Decide how to model custom fields

WooCommerce stores often contain custom product metadata for brand, material, compatibility, dimensions, seasonality, supplier data, or merchandising campaigns. Do not automatically create a Vespa field for every metadata key. First classify each value:

  • Searchable text: index it as text when users are likely to express it in a natural-language query.
  • Facet or filter: use a normalized string, numeric value, boolean, or array when users select it explicitly.
  • Ranking signal: use a numeric or date field when it affects ordering.
  • Display-only metadata: keep it summary-only or retrieve it from the source system.

A fixed schema is preferable for important fields because it makes query behavior, validation, and ranking predictable. For genuinely dynamic metadata, a structured representation can be appropriate, but the query and filtering strategy must be defined before ingestion. Avoid placing every custom field into one opaque JSON or string field if the application needs reliable numeric ranges or exact filters.

Make updates safe and repeatable

WooCommerce data changes frequently: price, stock, sale dates, ratings, and visibility can change without a product description changing. Separate ingestion logic into stable identity, content, taxonomy, commercial, and ranking updates where practical. Every update should be idempotent so that replaying a webhook or synchronization batch produces the same document state.

Use date_modified or another source revision marker to prevent an older synchronization job from overwriting newer data. When a product is deleted or made permanently unavailable, remove it from the appropriate Vespa document set rather than leaving an orphaned searchable record.

Before production rollout, test representative products including simple products, variable products, products with multiple attributes, products without prices, out-of-stock items, products with HTML-heavy descriptions, and products containing custom metadata. Verify that each field supports the exact search, filter, sort, and display behavior required by the storefront.

Trending posts
You might also like