WPSolr logo
  • WPSOLR
  • Conversion Tools
    • Parsing & Chunking
    • Search
      • Algolia  Search
      • Apache Solr  Search
      • Elasticsearch  Search
      • Google AI Commerce  Search
      • OpenSearch  Search
      • OpenSolr  Search
    • AI Search
      • OpenSearch  AI Search
      • PostgreSQL  AI Search
      • Vespa  AI Search
      • Weaviate  AI Search
    • AI Recommendations
      • Amazon Personalize  AI Recommendations
      • Recombee  AI Recommendations
    • AI Personalized Search & Recommendations
      • Algolia  AI Personalized Search & Recommendations
      • Google AI Commerce  AI Personalized Search & Recommendations
  • Learn
    • Documentation & Guides – Everything to get started fast
    • Youtube – 100 practical videos from setup to live site demos
    • Forums – Direct support from our team
  • Pricing
  • My licenses
  • Contact Us
  • WPSOLR
  • Conversion Tools
    • Parsing & Chunking
    • Search
      • Algolia  Search
      • Apache Solr  Search
      • Elasticsearch  Search
      • Google AI Commerce  Search
      • OpenSearch  Search
      • OpenSolr  Search
    • AI Search
      • OpenSearch  AI Search
      • PostgreSQL  AI Search
      • Vespa  AI Search
      • Weaviate  AI Search
    • AI Recommendations
      • Amazon Personalize  AI Recommendations
      • Recombee  AI Recommendations
    • AI Personalized Search & Recommendations
      • Algolia  AI Personalized Search & Recommendations
      • Google AI Commerce  AI Personalized Search & Recommendations
  • Learn
    • Documentation & Guides – Everything to get started fast
    • Youtube – 100 practical videos from setup to live site demos
    • Forums – Direct support from our team
  • Pricing
  • My licenses
  • Contact Us
  • WPSOLR
  • Conversion Tools
    • Parsing & Chunking
    • Search
      • Algolia  Search
      • Apache Solr  Search
      • Elasticsearch  Search
      • Google AI Commerce  Search
      • OpenSearch  Search
      • OpenSolr  Search
    • AI Search
      • OpenSearch  AI Search
      • PostgreSQL  AI Search
      • Vespa  AI Search
      • Weaviate  AI Search
    • AI Recommendations
      • Amazon Personalize  AI Recommendations
      • Recombee  AI Recommendations
    • AI Personalized Search & Recommendations
      • Algolia  AI Personalized Search & Recommendations
      • Google AI Commerce  AI Personalized Search & Recommendations
  • Learn
    • Documentation & Guides – Everything to get started fast
    • Youtube – 100 practical videos from setup to live site demos
    • Forums – Direct support from our team
  • Pricing
  • My licenses
  • Contact Us
Understand Where Your Store Is Losing Sales
Expand All Collapse All
  • What is WPSOLR ?
    • The standard WordPress SQL search
    • The WPSOLR search with Elasticsearch, Apache Solr, or Algolia
  • Your configuration journey, step by step
    • Install Apache Solr
    • Hosted Apache Solr and SolrCloud Services
    • Install Weaviate
    • Install Opensearch
    • Install Elasticsearch
    • Elasticsearch Hosting
    • Configure search
      • Getting started with search
      • Getting started with Ajax (live) search
      • Getting started with facets
      • 0. Connect your indexes
        • Create a Google Retail search index
        • Create an Elasticsearch index
          • Generate a test Elasticsearch index
          • Create an Elastic Elasticsearch index
          • Create an Elasticsearch index
          • Create a Qbox Elasticsearch index
          • Create an ElasticPress Elasticsearch index
          • Create an Aiven Elasticsearch or OpenSearch index
          • Create a Bonsai Elasticsearch index
          • Create an Amazon AWS Elasticsearch index
          • Create an ObjectRocket Elasticsearch index
          • Create a Cloudways Elasticsearch or Opensearch index
          • Create an Alibaba Cloud Elasticsearch index
          • Create a Compose Elasticsearch index
          • Connect to an Elasticsearch index
        • Create an Apache Solr index
          • Generate a test Apache Solr index
          • Create an Apache Solr index
          • Create a SearchStax SolrCloud index
          • Create an Opensolr Solr index
          • Connect to an Apache Solr index
        • Create an Opensearch index
          • Create a WPSOLR Hosting OpenSearch index
          • Create an Amazon AWS OpenSearch index
        • Create a Weaviate index
          • Weaviate with Google PaLM API
          • Weaviate with GPT4All
          • Weaviate with sentence transformers
          • Weaviate with CLIP text-to-image
          • Weaviate with HuggingFace Endpoints API
          • Weaviate with OpenAI GPT API
          • Weaviate with Cohere Multilingual API
          • Weaviate with Question Answering transformers
          • Weaviate with Hybrid search
          • Weaviate with Reranker - Cohere
          • Weaviate with Reranker - Transformers (cross-encoders)
        • Create an Algolia index
      • 1. Activate extensions (Add-ons)
        • bbPress add-on
        • YITH WooCommerce Ajax Search add-on
        • SEO add-ons
          • Yoast SEO add-on
          • All in One SEO add-on
        • Listable add-on
        • ACF add-on
        • Advanced Scoring add-on
        • Cron Scheduling add-on
        • Theme add-on
          • Filters layouts
            • Radiobox & Checkbox Layout
            • Numeric Range Layout
            • Colour Picker Layout
            • Date Picker (Flatpickr) Layout
            • Range Slider Layout
          • Add Ajax to the current Theme
          • Collapse taxonomy hierarchies
          • Custom Facets CSS
        • WPML add-on
        • Premium add-on
          • Manage more than one Elasticsearch or Solr index
        • PDF Embedder add-on
        • Geolocation add-on
        • AI Natural Language APIs add-on
          • Amazon Comprehend
          • Google Natural Language
          • Aylien Text Analysis
          • MeaningCloud
          • Qwam Text Analytics
        • Toolset Types add-on
        • AI Image and OCR APIs add-on
          • Google Vision
          • Amazon Rekognition
        • Embed Any Document add-on
        • MyListing add-on
        • Polylang add-on
        • WooCommerce add-on
        • Cross-domain federated search add-on
        • Directory+ add-on
        • Toolset Views add-on
        • Listify add-on
        • Jobify add-on
        • Query Monitor add-on
      • 2. Define your search
        • 2.3 Boosts
        • 2.3 Suggestions (Ajax search)
        • 2.1 Search
        • 2.2 Data
      • 3. Index your data
    • Configure recommendations
      • Algolia - Recommend and Personalized search
      • Recombee - Recommendations
      • Google Retail - AI Recommendations and Personalized search
      • Amazon Personalize
    • Configure WPSOLR Parsing and Chunking
    • Create your Custom Templates for Citations, Suggestions, Related Posts & Recommendations
  • Quick start

Configure WPSOLR Parsing and Chunking

3318 views 3 July 23, 2026 Updated on July 29, 2026

File Parsing and Chunking require a WPSOLR Enterprise license.

This documentation describe the configuration of your indexes for File Parsing and Chunking.

 

Table of Contents

Toggle
  • Parsing & Chunking connectors
    • Creating a WPSOLR Parsing & Chunking Connector
      • Configuration setting for LlamaParse
      • Configuration setting for Unstructured
    • Activate WPSOLR add-ons
      • Activate the WPSOLR Parsing and Chunking add-on
      • Activate the WPSOLR ACF add-on
    • Create two projects to configure the “Parsing and Chunking” and “Search” connectors
    • Assign the Parsing and Chunking connector to the project
    • Select the data to be chunked
  • Indexing and searching chunks
    • Create a WPSOLR Search index
    • Assign the search index to its own project
    • What are search citations?
    • Select the chunks to be indexed
  • WPSOLR Chunks metabox
    • What is the WPSOLR Chunks Metabox?
    • WPSOLR Chunks Metabox file states
    • Counting files chunks in local cache
    • Deleting local files in cache
  • Using Batches
  • Using Cron Scheduling
  • Showing Citations
    • Citations templates
      • Citations text templates
      • Citations Twig templates
    • Citations in search results
    • Citations in suggestions
  • Frequently Asked Questions
    • What is document chunking?
    • Why does WPSOLR use chunking?
    • Which document types can be chunked?
    • Which document parsing and chunking providers are supported?
    • How does WPSOLR decide where to split documents?
    • What is the difference between parsing and chunking?
    • Does WPSOLR store chunks as separate documents?
    • Can users search both original documents and chunks?
    • How does chunking improve AI search?
    • Does WPSOLR support semantic chunking?
    • What is the recommended chunk size?
    • Does chunking work with vector search?
    • Can chunking be combined with keyword search?
    • How do I configure WPSOLR Chunking?

Parsing & Chunking connectors

Creating a WPSOLR Parsing & Chunking Connector

WPSOLR uses Parsing & Chunking connectors to communicate with external document processing services such as LlamaParse or Unstructured.

A connector is responsible for:

  • Uploading files to the parsing service
  • Monitoring the processing status
  • Retrieving the extracted content
  • Returning the result to WPSOLR for chunk generation and indexing

This architecture makes it easy to support multiple parsing providers while keeping the same indexing workflow.

WPSOLR manages the indexing process and the connector to provide the extracted content.

 

Configuration setting for LlamaParse

  1. Create a free trial LlamaParse Cloud account and project
  2. Collect your LlamaParse’s project API key on your new project home page. You will paste it in WPSOLR settings later.
    Image wpsolr-llamaparse-api-key-1024x413.png of Configure WPSOLR Parsing and Chunking
  3. Create a Parse configuration in menu “Parse”. Save it.
    Image wpsolr-llamaparse-create-configuration-1024x515.png of Configure WPSOLR Parsing and Chunking
  4. Navigate to your LamaParse configuration by clicking on button “Configs” in menu “Parse”
    Image wpsolr-llamaparse-navigate-to-configurations-1024x243.png of Configure WPSOLR Parsing and Chunking
  5. Click on icon “View JSON” of your configuration, and copy the json settings.
    Image wpsolr-llamaparse-click-configuration-view-json-1024x234.png of Configure WPSOLR Parsing and Chunking
  6. Copy the JSON content. You will paste it in WPSOLR settings later.
    Image wpsolr-llamaparse-click-configuration-copy-json-scaled.png of Configure WPSOLR Parsing and Chunking
  7. Create a LlamaParse connector in WPSOLR. Copy your API key and JSON Parse settings in the connector fields.Create a new connector:
    Image wpsolr-new-connector.png of Configure WPSOLR Parsing and ChunkingCopy your API key and JSON settings, then save:
    Image wpsolr-connector-chunking-scaled.png of Configure WPSOLR Parsing and Chunking

    1. If you see an error, remove the extra parameter “product_type" before pasting JSON:
      Image wpsolr-llamaparse-json-settings-edit-error-parameter.png of Configure WPSOLR Parsing and Chunking

This is all you need to configure the LlamaParse connector. The remaining settings below are common to all parsing connectors, including LlamaParse and Unstructured.

Configuration setting for Unstructured

  1. Create a free trial Unstructured Cloud account
  2. Create your Unstructured API key. You will paste it in WPSOLR settings later.
    Image wpsolr-unstructured-api-key-scaled.png of Configure WPSOLR Parsing and Chunking
  3. Create a new Workflow configuration in menu “Workflow”. Save it.
    Image wpsolr-unstructured-new-workflow-277x300.png of Configure WPSOLR Parsing and ChunkingImage wpsolr-unstructured-new-workflow-custom-1024x622.png of Configure WPSOLR Parsing and Chunking
  4. Add a chunker transformer after the partitioner:
    Image wpsolr-unstructured-new-workflow-add-chunker-scaled.png of Configure WPSOLR Parsing and Chunking
  5. Configure the partitioner to your liking and budget:
    Image wpsolr-unstructured-new-workflow-setup-partitioner-scaled.png of Configure WPSOLR Parsing and Chunking
  6. Configure the chunker to your liking and budget:
    Image wpsolr-unstructured-new-workflow-setup-chunker-scaled.png of Configure WPSOLR Parsing and Chunking
  7. Save your workflow for reference. We will not use it, but its JSON configuration
  8. Click on menu “Show Options” of your Workflow
    Image wpsolr-unstructured-new-workflow-menu-show-option-scaled.png of Configure WPSOLR Parsing and Chunking
  9. Copy the JSON content of the “job_nodes” property. You will paste it in WPSOLR settings later.
    Image wpsolr-unstructured-new-workflow-menu-copy-json-scaled.png of Configure WPSOLR Parsing and Chunking
  10. Create a Unstructured connector in WPSOLR. Copy your API key and JSON Parse settings in the connector fields.
    Image wpsolr-connector-unstructured-scaled.png of Configure WPSOLR Parsing and Chunking

This is all you need to configure the Unstructured connector. The remaining settings below are common to all parsing connectors, including LlamaParse and Unstructured.

Activate WPSOLR add-ons

Activate the WPSOLR Parsing and Chunking add-on

The add-on will send chunks returned by the Parsing and Chunking connector to the search connector.

Image wpsolr-add-on-chunking-1024x286.png of Configure WPSOLR Parsing and Chunking

Activate the WPSOLR ACF add-on

The add-on will send all ACF file fields to the Parsing and Chunking connector.

Image wpsolr-add-on-acf-for-chunking-1024x307.png of Configure WPSOLR Parsing and Chunking

Create two projects to configure the “Parsing and Chunking” and “Search” connectors

You will need to setup the Parsing and Chunking connector, and a Search connector of your choice. Each connector requires its own project.

Click on link “Manages projects” in menu “2.1 Connector”:
Image wpsolr-chunking-create-project-link-1024x137.png of Configure WPSOLR Parsing and Chunking

Create two projects, one for the Parsing and Chunking connector, one for the Search connector:
Image wpsolr-chunking-create-projects-1024x566.png of Configure WPSOLR Parsing and Chunking

Assign the Parsing and Chunking connector to the project

Image wpsolr-chunking-assign-chunking-project-1024x281.png of Configure WPSOLR Parsing and Chunking

Select the data to be chunked

Image wpsolr-chunking-settings-data-scaled.png of Configure WPSOLR Parsing and Chunking

  • Expand shortcodes: expands shortcodes if you want to include chunking elements like image/audio/video galleries.
  • Embedded media: select the post type embedded media elements that you wish to chunk, like the featured image, embedded library images, or embedded library files.
  • Media type: select the media library attachment types that you wish to chunk and search.

 

Indexing and searching chunks

Create a WPSOLR Search index

For this documentation, we will use an local Opensearch index, but any other supported search engine would be fine. You can find detailed instructions here.

Image wpsolr-chunking-new-opensearch-connector-scaled.png of Configure WPSOLR Parsing and Chunking

Assign the search index to its own project

Image wpsolr-chunking-opensearch-search-settings-scaled.png of Configure WPSOLR Parsing and Chunking

  • “Results have no citations”: when the search index has not been configured with citations at all (see details below how to configure a search index with citations)
  • “Hide results citations”:  when the search index has been configured with chunks, but you want to filter out citations for this specific project
  • “Show results citations”:  when the search index has been configured with chunks, and you want to show citations for this specific project
    • “Maximum number of citations per file”:  let’s imagine 30 citations from the same pdf file were returned in results. You can use this parameter to show only a smaller number of citations for each file, like “1” in the above screen capture. With 50 citations from 12 files, only 12 citations will be shown in results.
    • “Maximum number of distinct files per post type”:  let’s imagine 70 citations from 10 distinct files were returned in results. You can use this parameter to show only a smaller number of files cited, like “5” in the above screen capture. # citation here: (5 files) x (1 citation per file ) =  5 citations in total instead of 70.

What are search citations?

Search citations are file chunks shown in results, inside their post type, with a position like the page number or the audio/video timestamps.

Select the chunks to be indexed

Image wpsolr-chunking-opensearch-data-settings.png of Configure WPSOLR Parsing and Chunking

  • Select the connectors from which to index chunks: check the Parsing & Chunking connector’s name(s) you created before. Notice that several connectors can appear in the list: a Parsing & Chunking connector can supply chunks to several search indexes, and a search connector can index chunks from several Parsing & Chunking connectors.
    • Use the Cron extension to schedule chunks indexing: recommended when you have lots of documents and file chunks.
    • Index downloaded chunks immediately : as soon as a file’s chunks are downloaded from the Parsing & Chunking connector’s Cloud provider, they are indexed and ready to be searched.

In our settings above, 3 post types are chunked and indexed: posts, pages, and attachments (“Media”). All post types will appear in search results and suggestions, enriched with their respective file citations. Media here is a post type attached to a media file. The same media file can also be attached to a post or page with embedded media, or media galleries, or ACF file fields. You can instead decide to show only standard post types like posts and pages and products, also enriched with their embeded file citations: just unselect “Media” in the above screen.

 

WPSOLR Chunks metabox

What is the WPSOLR Chunks Metabox?

The WPSOLR Chunks metabox appears on supported WordPress content after a document has been parsed and chunked.

It provides a clear view of the generated chunks, allowing you to inspect the text that will be indexed by your search engine.

This makes it easy to verify the parsing results, troubleshoot extraction issues, and understand exactly what content is available for AI-powered search, semantic search, and retrieval.

WPSOLR Chunks Metabox file states

  • Empty metabox: No file on this post type has been uploaded for chunking yet.
    Image wpsolr-chunking-metabox-empty-1024x502.png of Configure WPSOLR Parsing and Chunking
  • Metabox with files still processing or with errors: after publishing a post type, all files not yet in the chunking cache are uploaded to the Chunking Cloud provider. The file is marked as “Processing”, until it has completed or failed.
    Image wpsolr-chunking-metabox-processing-210x300.png of Configure WPSOLR Parsing and Chunking
    You can click on each file link to download everything returned by the Chunking Cloud provider.
    Below is an example of a file error, perfectly explained as being not supported by the selected Chunking Cloud provider:
    – Click on link “jfk.flac”:
    Image wpsolr-chunking-metabox-error-links-popup-229x300.png of Configure WPSOLR Parsing and Chunking
    – Then click on link “Download the chunking error”:
    Image wpsolr-chunking-metabox-error-content-popup-202x300.png of Configure WPSOLR Parsing and Chunking
  • Metabox with successful files with chunks: as soon as the file chunks are processed, WPSOLR download them in a local cache folder specific to the selected Parsing & Chunking connector
    Image wpsolr-chunking-file-cache-300x142.png of Configure WPSOLR Parsing and Chunking
    You can click on successful file link to download the chunks returned by the Chunking Cloud provider. You will see the file containing chunks stored in the local cache above.
    – Click on link “dropbox.pdf”:
    Image wpsolr-chunking-metabox-success-links-popup-191x300.png of Configure WPSOLR Parsing and Chunking
    – Then click on link “Download the chunks”:
    Image wpsolr-chunking-metabox-success-content-popup-300x298.png of Configure WPSOLR Parsing and Chunking

 

Counting files chunks in local cache

The metabox displays the total number of chunks in the connector’s local cache for the current edited post type. “6” chunks in the example below:

Image wpsolr-chunking-file-cache-chunks-count.png of Configure WPSOLR Parsing and Chunking

 

Deleting local files in cache

All files being cached locally, including errors and processing files, if you want to restart the chunking process you just have to delete the cache and publish the post type.

– Select one or several files, click on “Delete file chunks and raw response” and publish the post type again:
Image wpsolr-chunking-file-cache-delete-226x300.png of Configure WPSOLR Parsing and ChunkingImage wpsolr-chunking-file-cache-deleted-243x300.png of Configure WPSOLR Parsing and Chunking

 

Using Batches

  1. Select your chunking connector to manually start the chunking process:
    Image wpsolr-chunking-operations-batch-chunking-1-scaled.png of Configure WPSOLR Parsing and Chunking

    • “Files uploaded for chunking”
      These are the post types with chunks to upload to the Cloud Chunking API.
    • “File chunks downloaded”
      After uploading files, WPSOLR will fetch the Cloud Chunking API to download their chunks. You may have to try several times until all chunks are processes or in error.
    • “File chunks with errors”
      WPSOLR detected an error returned by the Cloud Chunking API. You can follow the error links to filter the post types with errors, then edit each post type to check errors in their metabox:
      Image wpsolr-chunking-chunking-filter-errors-scaled.png of Configure WPSOLR Parsing and Chunking
      Image wpsolr-chunking-metabox-error-content-popup-202x300.png of Configure WPSOLR Parsing and Chunking
  2. Select your search index connector to manually start the indexing process:
    Image wpsolr-chunking-operations-batch-chunking-scaled.png of Configure WPSOLR Parsing and Chunking

    • “Documents to process”
      Index all post types with or without existing chunks.  Chunks should have been processed first as described above.
    • “Chunked files to process”
      WPSOLR detected some post types with new chunks available for indexing.

 

Using Cron Scheduling

Considering the asynchronous nature of the Cloud Chunking API, using a cron to schedule everything makes sense.

  1. Activate the WPSOLR Cron Scheduling add-on
  2. Set the chunking connector in first position (use drag&drop if necessary)
    Image wpsolr-chunking-chunking-cron.png of Configure WPSOLR Parsing and Chunking
  3. Set the search connector in second position
    Image wpsolr-chunking-indexing-cron.png of Configure WPSOLR Parsing and Chunking

You can start a cron every night, every hour or more frequently. You can also schedule chunking and indexing independently: chunking more frequently than indexing for instance.

Showing Citations

Citations templates

Citations text templates

You can modify the citations text templates in menu “2.6 Texts”. Those texts can also be translated with WPSOLR WPML or Polylang add-ons.

Image wpsolr-chunking-texts-settings-1024x413.png of Configure WPSOLR Parsing and Chunking

Citations Twig templates

For more advanced modifications, you can duplicate the WPSOLR Twig templates into your active theme’s directory.

Citations in search results

  • Chose a Twig citations template on the search connector:

Image wpsolr-chunking-suggestions-template-settings-scaled.png of Configure WPSOLR Parsing and Chunking

  • The citations will appear on the active theme’s search page results managed by this search connector:

Image wpsolr-chunking-search-template-shown-736x1024.png of Configure WPSOLR Parsing and Chunking

 

Citations in suggestions

  • Chose a citations Twig template on the search connector’s suggestions settings :
    Image wpsolr-chunking-suggestions-template-settings-1-scaled.png of Configure WPSOLR Parsing and Chunking
  • The citations will appear on the suggestions managed by this search connector:
    Image wpsolr-chunking-suggestions-template-shown-198x300.png of Configure WPSOLR Parsing and Chunking

Frequently Asked Questions

What is document chunking?

Document chunking is the process of splitting large documents into smaller, meaningful sections called chunks. Each chunk can then be independently indexed, searched, and used by search applications.

Instead of treating a complete document as a single block of text, WPSOLR Chunking creates smaller searchable units that improve relevance and accuracy.

Why does WPSOLR use chunking?

Large documents often contain multiple topics. Indexing them as a single document can make search results less precise because the matching content may represent only a small part of the document.

Chunking allows WPSOLR to:

  • Return more relevant search results.
  • Improve AI-generated answers by retrieving only the most useful passages.
  • Handle large documents that exceed language model context limits.
  • Preserve relationships between document sections.

Which document types can be chunked?

WPSOLR Chunking can process many types of content, including:

  • PDF documents.
  • Office documents.
  • Text files.
  • HTML content.
  • WordPress posts and pages.
  • Other formats supported by connected parsing services.

The available formats depend on the configured parser connector.

Which document parsing and chunking providers are supported?

WPSOLR supports leading document processing providers to extract and structure content before indexing.

You can choose the solution that best matches your requirements:

  • LlamaParse — Advanced document parsing designed to extract high-quality structured content from complex documents such as PDFs, reports, and technical documentation.
  • Unstructured — A flexible document processing platform supporting a wide range of file formats and extraction workflows.

Both providers integrate with the WPSOLR Chunking workflow, allowing you to transform documents into optimized searchable chunks for traditional search, vector search, and AI-powered retrieval.

You keep control of your content processing strategy while WPSOLR provides the search and retrieval infrastructure.

 

How does WPSOLR decide where to split documents?

WPSOLR can use different chunking strategies depending on the document structure and configuration.

Chunking can be based on:

  • Paragraph boundaries.
  • Document sections and headings.
  • Sentence boundaries.
  • Token or character limits.
  • Semantic segmentation provided by AI-powered parsers.

What is the difference between parsing and chunking?

Parsing extracts the content and structure of a document.

Chunking divides the extracted content into smaller searchable pieces.

For example:

  • A PDF parser extracts text, headings, tables, and metadata.
  • The chunking process splits this extracted content into optimized sections for search and AI retrieval.

Does WPSOLR store chunks as separate documents?

Yes. WPSOLR can index chunks as individual searchable records while keeping references to the original document.

Each chunk can contain:

  • The extracted text.
  • The parent document identifier.
  • Chunk position information.
  • Metadata from the original document.
  • Additional fields required for filtering and ranking.

Can users search both original documents and chunks?

Yes. WPSOLR can be configured to search chunk-level content while maintaining links back to the original document.

This allows users to find the exact relevant section while still accessing the complete source document.

How does chunking improve AI search?

AI search systems retrieve information before generating an answer. If the retrieved content is too large or contains unrelated information, the generated answer quality decreases.

Chunking improves AI search by:

  • Reducing irrelevant context.
  • Increasing retrieval precision.
  • Providing focused passages to language models.
  • Supporting vector and hybrid search.

Does WPSOLR support semantic chunking?

Yes. When using compatible AI parsing or indexing connectors, WPSOLR can support semantic approaches where chunks are created based on meaning rather than only fixed sizes.

Semantic chunking helps preserve complete ideas and improves retrieval quality for complex documents.

What is the recommended chunk size?

There is no universal ideal chunk size. It depends on:

  • Document type.
  • Search engine.
  • Embedding model.
  • Language model context size.
  • Expected user queries.

WPSOLR allows administrators to adjust chunking settings to optimize results for their specific use case.

Does chunking work with vector search?

Yes. Chunking is especially useful for vector search because embeddings work best when each vector represents a focused piece of information.

Instead of generating one embedding for an entire document, WPSOLR can generate embeddings for individual chunks, improving similarity matching.

Can chunking be combined with keyword search?

Yes. WPSOLR supports combining chunk-based retrieval with traditional keyword search.

Hybrid search can combine:

  • Exact keyword matching.
  • Full-text relevance scoring.
  • Vector similarity.
  • Metadata filtering.

This provides better results across both traditional and AI-powered search experiences.

How do I configure WPSOLR Chunking?

Configuration depends on the selected parsing and indexing workflow.

Typical steps are:

  1. Enable document parsing.
  2. Configure a parser connector.
  3. Select chunking options.
  4. Index the generated chunks.
  5. Test search and AI retrieval quality.

Was this helpful?

3 Yes  No
Related Articles
  • Create a WPSOLR Hosting OpenSearch index
  • Create your Custom Templates for Citations, Suggestions, Related Posts & Recommendations
  • Getting started with facets
  • Getting started with Ajax (live) search
  • Getting started with search
  • 3. Index your data

Didn't find your answer? Contact Us

Previously
Amazon Personalize
Up Next
Create your Custom Templates for Citations, Suggestions, Related Posts & Recommendations

Recommended content

Powered by WPSOLR Enterprise with Recombee

Understand Where Your Store Is Losing Sales
Join our Affiliate Program!
Features
Search Hosting for WordPress
File Parsing and Chunking
Integration Hub
Search
Facets (Filters)
Search Add-ons
Recommendations
Keyword search engines
Comparison grid
Apache Solr & Solr Cloud
Elasticsearch
OpenSearch
Algolia
AI search engines
Comparison grid
PostgreSQL pgvector
Weaviate
Google Retail
Vespa.ai
OpenSearch vector
AI recommendation engines
Comparison grid
Recombee
Amazon Personalize
AI personalization engines
Comparison grid
Google Retail
Algolia
Buy and Learn
Pricing
Privacy policy
Terms and Conditions
© 2026 Eostis SARL. All Rights Reserved.
More
My licenses
Support
Documentation
Blog
Recommendations powered for you by WPSolr
Manage Consent
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.
  • Manage options
  • Manage services
  • Manage {vendor_count} vendors
  • Read more about these purposes
View preferences
  • {title}
  • {title}
  • {title}