Table of contents :

Product Titles, Descriptions and Attributes in Vespa

100-days-of-vespa-ai-woocommerce-006

Table of contents :

Model the Product Data Before You Rank It

WooCommerce product data usually arrives as a combination of post fields, taxonomy terms, custom attributes, and variation records. Vespa works best when those values are mapped deliberately into a document schema rather than copied into one large text field.

A practical product document might contain:

  • product_id as the stable WooCommerce product identifier.
  • title for the customer-facing product name.
  • description for the long description and relevant short-description content.
  • brand, category, and attribute_* fields for filtering and structured matching.
  • price, stock_status, and visibility for commerce constraints.
  • popularity, sales signals, or business rules used during ranking.

Keep identifiers and filterable values separate from prose. A title such as “Red Cotton T-Shirt” is useful for full-text search, while the color and material should also be represented as structured values when shoppers need to filter by them.

Define Text Fields for Titles and Descriptions

In Vespa, a field can be indexed for full-text search or stored as an attribute for fast access and filtering. Product titles and descriptions normally need indexing. A simplified schema can look like this:

schema product {
    document product {
        field product_id type string {
            indexing: summary | attribute
        }
        field title type string {
            indexing: index | summary
            index: enable-bm25
        }
        field description type string {
            indexing: index | summary
            index: enable-bm25
        }
        field brand type string {
            indexing: index | summary | attribute
        }
        field category type string {
            indexing: index | summary | attribute
        }
        field color type string {
            indexing: index | summary | attribute
        }
        field material type string {
            indexing: index | summary | attribute
        }
    }
}

summary makes a field available in document responses. index makes text searchable. attribute is appropriate when the value must be used efficiently for filtering, sorting, grouping, or ranking features. A field can use more than one indexing mode when the application needs both behaviors.

Use a stable, normalized representation for attributes. For example, map “navy”, “dark blue”, and “blue navy” to the value your catalog uses consistently. Normalization should happen during the WooCommerce-to-Vespa feed process, not separately in every query.

Titles Need More Weight Than Long Descriptions

For product search, a match in the title is often a stronger signal than the same term in a long description. Vespa’s BM25 rank feature can score indexed text, and a ranking profile can combine title and description relevance explicitly.

rank-profile product_text inherits default {
    first-phase {
        expression: 2.5 * bm25(title) + bm25(description)
    }
}

The exact weight should be measured with search judgments or production metrics. A title multiplier of 2.5 is only an example. If the catalog contains short, repetitive titles, over-weighting them can make results less useful. If descriptions contain extensive marketing copy, giving the title more influence can prevent weak description matches from outranking highly relevant products.

For fields with different importance, separate indexed fields are preferable to concatenating title, description, and attributes into one string. Separate fields allow ranking, highlighting, and query interpretation to evolve independently.

Clean WooCommerce Content Before Feeding Vespa

WooCommerce descriptions frequently contain HTML, shortcodes, tracking fragments, repeated headings, and boilerplate shipping text. Strip presentation markup while preserving meaningful text. Remove repeated content that appears on many products because it contributes little to relevance and can create noisy matches.

For example, this source content:

<h2>Organic Cotton Hoodie</h2>
<p>Soft 100% cotton hoodie with a brushed interior.</p>
[shipping_note]Free delivery on orders over $50[/shipping_note]

could become a title of Organic Cotton Hoodie and a description of Soft 100% cotton hoodie with a brushed interior. Shipping information should normally be modeled as a business or fulfillment field rather than mixed into product relevance text.

Preserve the original product identifier and a canonical URL. If a product is updated or deleted in WooCommerce, the feed must update or remove the corresponding Vespa document so stale titles and descriptions do not remain searchable.

Represent Attributes for Search and Filters

Attributes serve several different purposes:

  • Exact filtering: color, size, gender, material, and compatibility values.
  • Faceting: counts such as the number of products available in each size or brand.
  • Text matching: attributes that shoppers may mention naturally in a query.
  • Display: values shown in the result card or product detail page.

Do not assume that one representation is ideal for all four purposes. A color can be stored as an attribute for filtering and also included in an indexed text field if queries such as “blue running shoes” must match it. A multi-valued attribute, such as compatible device models, should be represented as a collection when a product can have several values.

For a small, stable catalog, explicit fields such as color, size, and material are easy to query and maintain. For a catalog with many dynamic WooCommerce attributes, a structured array or mapped representation may reduce schema changes, but it requires a clear convention for names, normalization, filtering, and response formatting. Choose based on catalog stability and query requirements rather than trying to mirror WordPress metadata exactly.

Map a WooCommerce Product into a Vespa Document

A product feed should transform the WooCommerce response into the Vespa document format used by the application. For example, a product might be represented conceptually as:

{
  "put": "id:catalog:product::1842",
  "fields": {
    "product_id": "1842",
    "title": "Organic Cotton Hoodie",
    "description": "Soft 100% cotton hoodie with a brushed interior.",
    "brand": "Northwind",
    "category": "Hoodies",
    "color": "forest green",
    "material": "organic cotton",
    "stock_status": "instock",
    "price": 64.00
  }
}

The document ID should remain stable when the product title or description changes. Stable IDs make partial updates and deletion handling predictable. Variation products require an explicit decision: index each variation separately when size, color, price, or inventory affects search results, or index the parent product with aggregated variation data when shoppers should land on one parent product page.

Query Titles, Descriptions, and Attributes Together

A basic query can search the indexed product fields using Vespa’s query language. A user searching for “green organic cotton hoodie” should be able to match the title, description, and structured attribute values.

select * from product where userQuery();

The application can send the user’s text as the query input and select the ranking profile:

ranking=product_text

For explicit filters, add structured constraints rather than placing filter terms into free text. For example, a color filter should be represented as a condition on the normalized color field, while the phrase “lightweight hoodie” remains a relevance query. This separation improves facet counts, ranking behavior, and query debugging.

When a shopper selects several values, decide whether the values are alternatives or requirements. “Red or blue” is an OR condition; “cotton and size medium” is an AND combination across fields. The WooCommerce storefront should generate these constraints consistently and validate allowed values against the catalog.

Use Search Summaries for Result Cards

Search results should return only the fields required by the storefront. A document summary can expose the title, short description, product URL, image URL, price, and selected attributes without returning the complete source document.

document-summary product-result {
    summary title {}
    summary description {}
    summary product_id {}
    summary brand {}
    summary category {}
    summary color {}
    summary material {}
    summary price {}
}

Keep result text suitable for the WooCommerce card. A long description may be useful for matching but unsuitable for display. Many stores therefore maintain a concise short_description field for summaries while retaining a cleaned full description field for search and product detail pages.

Test Relevance with Real Catalog Queries

Build an evaluation set from actual WooCommerce searches and manually label useful results. Include exact product names, attribute-heavy queries, category queries, misspellings, and searches that should return no products. Compare title-only ranking, title-plus-description ranking, and title-plus-attribute ranking.

Watch for common failures:

  • A product with a matching boilerplate description outranks a product with the requested term in its title.
  • Different spellings of an attribute produce separate facets.
  • Variation inventory is ignored even though the parent product appears available.
  • HTML or shortcode text creates irrelevant matches.
  • Deleted or unpublished WooCommerce products remain in Vespa.

Use query logs, click-through data, add-to-cart events, and conversion signals carefully. These signals can support ranking, but they should not override hard constraints such as catalog visibility, stock policy, price range, or customer-selected attributes.

Trending posts
You might also like