Skip to main content

Overview

From the Search Settings page, you can configure the embedding model, reranking, and a variety of advanced search and indexing options. Search settings overview page

Embedding Model

The embedding model is used to convert your documents into vectors that are stored in OpenSearch. These vectors are used to search for relevant documents when a user queries Onyx. A powerful embedding model can significantly improve the accuracy of your search results, but comes at the cost of additional memory and disk usage. Embedding model configuration page
The best embedding models are generally available through a cloud provider such as Cohere or Google.To use a cloud provider, select the model you want to use, configure its authentication, and click Connect. Then click Apply & Re-index to use the new model.
Self-hosted models run on your own infrastructure and guarantee that your data does not leave your bounds.
Embedding data at the scale that Onyx operates requires significant compute resources. If you want to use a self-hosted model, we strongly recommend you make a GPU available to Onyx’s indexing model server container.
In the Self-hosted tab, you can select a suggested model or follow the instructions to connect your own model.

Google Embeddings

Google embedding models use Vertex AI. In Index Settings, open View All Models → Cloud-based and click Connect on a Google model. Google supports two authentication methods:
  • Service Account JSON: upload the service account’s JSON key. Existing Google embedding configurations continue to use this method.
  • Workload Identity (GKE): available on self-hosted Onyx. Onyx uses Application Default Credentials from the pods that make embedding requests, so no JSON key is required.
Google embedding authentication menu with Service Account JSON and Workload Identity (GKE) options For Workload Identity, follow the GKE deployment and IAM setup. Grant roles/aiplatform.user in the Vertex AI project to every identity used by the API server and embedding workers, including document-processing and user-file processing workers. Those pods must run on node pools with the GKE Metadata Server enabled. Cloud embeddings are called from these services, rather than the self-hosted model server. In the Google connection form, select Workload Identity (GKE) and enter the GCP Project ID where Vertex AI is enabled. This project can differ from the GKE cluster’s project. Set Google Cloud Region Name to a location supported by your embedding model; the default is global. The Vertex AI location is independent of the cluster’s location. Google embedding form with Workload Identity (GKE), an example GCP Project ID, and the global region Click Connect to test the configuration with a real embedding request before saving. Switching from Service Account JSON to Workload Identity removes the stored JSON key. You can later change the authentication method, project, or region using Google’s Edit credentials button. After selecting the model, click Apply & Re-index. Verify that indexing completes and a search retrieves the expected documents. A successful connection test verifies the API server’s credentials; indexing also requires the workers to have Vertex AI access. If the connection or indexing fails, check the pod’s Kubernetes service account, its node pool’s metadata server, and the IAM grant in the configured Vertex AI project. For a linked Google Cloud service account, also check the iam.gke.io/gcp-service-account annotation and roles/iam.workloadIdentityUser binding. Ensure the Vertex AI API is enabled and allow time for new IAM grants to propagate. Do not set GOOGLE_APPLICATION_CREDENTIALS to a JSON key file when you intend to use the GKE pod identity: ADC checks that variable before the metadata server.

Embedding Swaps

If you select a new embedding model, Onyx will need to re-index all of your data. During this process, the old embedding model will still be available for searches. While the swap is in progress, you will see the Search Settings page show details indexing progress.
This process can take a while. Additionally, private user data is also being re-indexed, but are not displayed to the Admin Search Settings page.
Embedding swap in progress

Vector Quantization

Vector quantization stores the document vectors in OpenSearch with fewer bits for each dimension. This decreases the memory that the vector index uses. Fewer bits can decrease the search quality. To set it, go to Index Settings and use the Vector Quantization field in the Embedding Model section. When the index is quantized, Onyx rescores the top search results with the full-precision vectors. OpenSearch keeps the full-precision vectors on disk for this, so the disk usage increases slightly.
The quantization is part of the index. When you change it, Onyx re-indexes all of your data into a new index, the same as an embedding swap.
The OpenSearch that ships with Onyx (Docker Compose and Helm) is version 3.6.0. If you connect Onyx to an external OpenSearch cluster that is older than 3.6, Onyx does not accept the 1-bit option. This setting is not available on Onyx Cloud.
Vector Quantization options on the Index Settings page

Reranking

Reranking is an optional step that can be used to improve the accuracy of your search results. A reranking model will assess and re-organize your search results based on the relevance of the documents to the query. This process adds a small amount of latency to your search results. Generally, re-ranking is only useful if you have a very large number of documents. Reranking configuration page

Advanced Configs

On the final page, you can configure a variety of advanced search settings.
Multilingual expansion rephrases your queries into the specified other languages. This can be helpful for cross-language results.
Multipass indexing creates chunks of varying sizes and stores them in the index. This can help the hybrid search algorithm better identify relevant sources.
Contextual RAG adds additional document-level information to every chunk in the index. This can help the hybrid search algorithm better identify relevant sources.
Contextual RAG can be very expensive as it adds a signficnat amount of data to every embedding call.
Setting this value will reduce the number of dimensions in the embedding vectors. This can reduce the memory usage of the index, but may reduce the accuracy of the results.
Reduced dimension is only supported for OpenAI embedding models at this time.
Advanced search configurations page