Performance guide
This guide describes application-side methods that reduce latency when you use Zep.
Retrieval latency depends on the graph size, query, filters, result limits, deployment, and application network path. Measure end-to-end latency with representative data and requests. The methods below reduce avoidable application-side latency.
Reuse the Zep SDK Client
The Zep SDK client maintains an HTTP connection pool. Reuse the client to avoid creating a connection for each request:
- Create a single client instance and reuse it across your application
- Avoid creating new client instances for each request or function
- Consider implementing a client singleton pattern in your application
- For serverless environments, initialize the client outside the handler function
Optimizing Context Operations
The thread.add_messages and thread.get_user_context methods are optimized for conversational messages and low-latency retrieval. For optimal performance:
- Use
graph.addfor larger documents, tool outputs, or business data (up to 10,000 characters per call) - Chunk large documents before adding them to the graph
- Remove unnecessary metadata or content before persistence
- For bulk document ingestion, use
zep-ingestor the Batch API rather than parallelgraph.addcalls
Get the Context Block sooner
You can request the Context Block directly in the response to the thread.add_messages() call.
This optimization eliminates the need for a separate thread.get_user_context() call.
Read more about our Context Block.
In this scenario you can pass in the return_context=True flag to the thread.add_messages() method.
Zep will perform a user graph search right after persisting the data and return the context relevant to the recently added messages.
Searching the Graph Sooner
Use graph search and Advanced Context Block construction when you need custom parameters. You can search the graph while another request adds data.
Construct a custom Context Block from the search results. Advanced Context Block construction provides examples.
Optimizing Search Queries
Zep uses hybrid retrieval combining semantic (vector) similarity, BM25 full-text search, and graph traversal in a single ranked result. For optimal performance:
- Keep your queries concise. The public graph-search API accepts a maximum of 400 characters and rejects a longer query.
- Longer queries may not improve search quality and will increase latency
- Consider breaking down complex searches into smaller, focused queries
- Use specific, contextual queries rather than generic ones
Best practices for search:
- Keep search queries concise and specific
- Structure queries to target relevant information
- Use natural language queries for better semantic matching
- Consider the scope of your search (graphs versus user graphs)
Warming the User Cache
Use user.warm before a likely retrieval to prepare the user’s graph data. For example, call it when a user logs in to your service or opens your application.
Summary
- Reuse Zep SDK client instances to optimize connection management
- Use appropriate methods for different types of content (
thread.add_messagesfor conversations,graph.addfor large documents) - Keep search queries focused and under the token limit for optimal performance
- Warm the user cache when users log in or open your app for faster retrieval