<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:sy="http://purl.org/rss/1.0/modules/syndication/" version="2.0">
  <!-- Source: https://gomomento.com/feed -->
  <channel>
    <title>Momento</title>
    <atom:link href="https://siftrss.com/f/0BgwYvPmQz" rel="self" type="application/rss+xml"/>
    <link>https://siftrss.com/f/0BgwYvPmQz</link>
    <description>An enterprise-ready serverless platform for caching and pub/sub</description>
    <lastBuildDate>Tue, 06 Oct 2026 18:00:00 GMT</lastBuildDate>
    <language>en</language>
    <sy:updatePeriod>hourly</sy:updatePeriod>
    <sy:updateFrequency>1</sy:updateFrequency>
    <image>
      <url>https://www.gomomento.com/wp-content/uploads/2024/06/cropped-favicon-green-32x32.png</url>
      <title>Momento</title>
      <link>https://www.gomomento.com/</link>
      <width>32</width>
      <height>32</height>
    </image>
    <item>
      <title>Beacon ingestion with Valkey Router and Momento Cache</title>
      <link>https://www.gomomento.com/use-cases/beacon-ingestion-with-valkey-router/</link>
      <pubDate>Tue, 06 Oct 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <category><![CDATA[Media]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/use-cases/beacon-ingestion-with-valkey-router/</guid>
      <description><![CDATA[<p>Protect live-event telemetry ingestion with Valkey Router: batch small beacons, group related events and buffer bursts before downstream processing.</p>]]></description>
      <content:encoded><![CDATA[<p>Live events concentrate activity across large populations of TVs, browsers and mobile devices. Those devices send beacons as people watch, interact and experience changes in playback. The resulting traffic can surge suddenly, and its profile is difficult to predict. Critical internal services need to keep processing event data while the audience changes the volume and pace of ingestion.</p>
<p>A beacon carries device telemetry with selected signals for downstream analysis. The event data it carries might support <a href="https://www.gomomento.com/blog/fox-monitors-super-bowl-viewership-experience-with-real-time-data-insights/">quality-of-experience monitoring</a>, <a href="https://www.gomomento.com/blog/stop-cdn-leeching-with-concurrency-tracking/">anti-piracy detection</a> or <a href="https://www.gomomento.com/blog/beyond-the-goals-three-ways-momento-scales-the-football-world-cup-in-real-time/">personalization</a>. Selecting those signals matters to ingestion as well as analysis. Every added field increases the payload, and those extra bytes multiply across the beacon traffic that the system must receive and hold.</p>
<p>Beacon ingestion therefore needs a protective boundary between device traffic and internal processing. In this architecture, <a href="https://www.gomomento.com/products/valkey-router/">Momento Valkey Router</a> receives beacons through a gateway tier, manages device connections and batches data for forwarding. Shared, open-source Valkey collects related events and buffers them until downstream services can process them. Together, the gateway and Valkey turn an unpredictable beacon firehose into organized data that internal services can consume at a manageable rate.</p>
<p><img src="https://www.gomomento.com/assets/content/use-cases/beacon-ingestion-with-valkey-router/architecture.png" alt="TV, web and mobile clients send beacons along three paths into a logical Valkey Router tier and shared Valkey; three separate event-stream arrows lead to boxed Kafka, ClickHouse and S3 icons representing downstream choices"></p>
<p><em>Valkey Router and shared Valkey collection form a protective ingestion boundary between device traffic and downstream services. Kafka, ClickHouse and S3 illustrate downstream choices.</em></p>
<h2 id="protect-internal-services-from-device-traffic">Protect internal services from device traffic</h2>
<p>The gateway separates the audience’s connections from the connections that internal infrastructure must support. Without that boundary, a growing device population can increase <a href="https://www.gomomento.com/blog/understanding-the-nxm-problem-in-distributed-caches/">database connection pressure</a> as well as beacon volume. These are distinct demands. An internal service might have enough processing capacity for the events while struggling to manage the connections delivering them.</p>
<p>Valkey Router handles device connections and <a href="https://www.gomomento.com/blog/why-large-cache-systems-need-routing-layers/">reuses a pool of connections to Valkey</a>. Pooling prevents each device connection from becoming a separate Valkey connection. The gateway can receive traffic from many devices while controlling the number of database connections. Momento supplies this gateway layer alongside Valkey’s open-source storage foundation, placing device-facing capabilities ahead of shared event collection.</p>
<p>The gateway also controls access to that collection. Authentication and access checks determine whether incoming traffic is allowed to reach internal infrastructure. HTTPS and TLS termination protect device-to-gateway communication, and custom certificates let the endpoint use the certificates required by the deployment. These capabilities put connection and access handling at the boundary where device traffic enters the system.</p>
<p><a href="https://www.gomomento.com/blog/did-you-say-you-want-a-distributed-rate-limiter/">Throttling</a> limits the forwarding rate when incoming traffic would otherwise overwhelm the next stage. The gateway controls ingestion before excess traffic reaches the database and consumers. Connection pooling and throttling address different parts of the same protection. One controls database connections, while the other controls the rate of data reaching internal services.</p>
<h2 id="share-forwarding-overhead-across-small-beacons">Share forwarding overhead across small beacons</h2>
<p>Managing connections still leaves a large number of small payloads to forward. Sending every beacon separately repeats communication and request-handling overhead. When that repeated overhead consumes a substantial part of ingestion capacity, the system spends effort moving small pieces of telemetry instead of collecting their event data efficiently.</p>
<p>Valkey Router batches compatible beacons locally before forwarding them. Several beacons share the overhead of a batch, reducing the repeated forwarding cost per beacon. Their event data remains available for collection. Batching changes how data is transferred without requiring the application to remove the observations it needs for analysis.</p>
<p>Batching introduces a short wait at the gateway because data must accumulate before it can be sent together. The gateway bounds that wait and the amount of data it gathers. A count trigger sends a batch when enough beacons have accumulated. A timer sends the accumulated data when the waiting interval ends. During busy periods, the count trigger can fill batches quickly. During quieter periods, the timer lets small batches progress instead of leaving beacons waiting for traffic that may not arrive soon.</p>
<p>A size bound can also limit how much data a batch contains. This matters because beacon sizes can vary with the selected signals. A batch with an acceptable beacon count could still contain too many bytes if those beacons have <a href="https://www.gomomento.com/blog/why-large-payloads-break-caches-at-scale/">larger payloads</a>. Count, time and size bounds serve complementary purposes. They limit local accumulation while balancing forwarding efficiency against the delay introduced by waiting.</p>
<p>Each gateway assembles its own batches from the traffic it receives. This keeps batching close to device ingestion, where small payloads first accumulate. The local batch holds data briefly to improve transfer efficiency. Shared Valkey collection serves the longer-lived need to hold events for downstream processing. Efficient batches from separate gateways must also reach the right collection point, because each gateway may receive only part of a partition’s activity.</p>
<h2 id="bring-related-events-to-the-same-collection-point">Bring related events to the same collection point</h2>
<p>A gateway tier distributes ingestion across multiple instances, but related beacons can reach different gateways. A consumer processing an application-defined partition needs its related events together, even when they come from different devices. Leaving those events fragmented across gateway-local collections would push the task of gathering them onto every downstream consumer.</p>
<p>The architecture collates related events by deriving a grouping key deterministically from beacon properties. Every gateway applies the same grouping rule. Beacons with the same grouping property produce the same grouping key, giving their related activity a consistent logical address. The gateways use that address to route event data to the same <a href="https://www.gomomento.com/blog/horizontal-scaling-with-elasticache-redis-stop-getting-burned-by-hot-keys-and-shards/">Valkey shard</a>. The shared collection brings related events together in a coherent data stream for downstream consumption.</p>
<p>Consider two devices sending beacons assigned to partition <code>S42</code>. One beacon reaches one Valkey Router instance and the other reaches another. Both carry the partition identifier, and both gateways derive the grouping key <code>partition:S42</code> from it. Each gateway uses that logical address to route the related event data to the selected Valkey shard. The partition’s events meet in shared collection even though different devices and gateways handled the beacons.</p>
<p>The selected shard is a portion of shared Valkey collection that holds multiple keyed groups. The grouping key keeps the partition’s events associated alongside data for other partitions. A consumer can process the collected partition activity without gathering fragments from individual gateways. This separates the gateway that receives a beacon from the collection where its related events belong, allowing the ingestion tier to distribute device traffic while preserving the relationships downstream analysis needs.</p>
<p><img src="https://www.gomomento.com/assets/content/use-cases/beacon-ingestion-with-valkey-router/batching.png" alt="TV, web (globe) and mobile clients send distinct S42 events marked by a dot, an x and a plus through three independent Valkey Routers. Mixed local buffers feed event batches; the three tracked S42 events meet with one other S42 event in partition:S42 in the middle Valkey shard. Muted shards above and below hold S41 triangles and S43 diamonds."></p>
<p><em>Each gateway buffers events locally for efficient forwarding in batches. Matching partition keys bring related events from different gateways to the same Valkey collection point. The arrows follow the three marked S42 events; other partitions and shard placements are illustrative.</em></p>
<h2 id="give-events-room-to-wait-for-consumers">Give events room to wait for consumers</h2>
<p>Organized events can still arrive faster than downstream services process them. A traffic burst increases ingestion immediately, while a consumer may continue at its previous processing rate. Consumer slowdowns can create the same imbalance even when incoming traffic remains steady. Shared buffering gives the architecture somewhere to hold that difference.</p>
<p>Valkey holds collected event data while consumers process it. During a burst, the buffer grows as ingestion outpaces consumption. When traffic subsides or consumer capacity increases, consumers can catch up and the backlog can shrink. The shared buffer lets the gateway continue receiving traffic through a temporary mismatch in rates, within the capacity available to hold the data.</p>
<p>That capacity is finite. The amount of data waiting depends on both incoming volume and how long it waits for downstream processing. Larger beacon payloads or heavier traffic increase the data entering the buffer. Slower consumers extend the waiting duration. Either change can increase the space required, and both can occur during the same live event.</p>
<p>Headroom accommodates those departures from normal operation. It provides space for traffic spikes, periods of slower consumption and the interval before autoscaling adds capacity. <a href="https://www.gomomento.com/blog/best-practices-for-elasticache-redis-autoscaling-and-how-to-do-better/">Autoscaling takes time to respond to increased demand</a>. Existing capacity must hold the accumulating data during that response interval, so scaling complements headroom rather than replacing it.</p>
<p>This buffer serves a different purpose from gateway batching. The gateway waits briefly to forward small beacons efficiently. Valkey holds shared, collected events while downstream processing catches up. Keeping those responsibilities distinct explains why efficient forwarding alone cannot absorb a sustained difference between ingestion and consumption.</p>
<h2 id="hand-collected-events-to-downstream-services">Hand collected events to downstream services</h2>
<p>At the downstream boundary, internal services receive event data that has been collected, grouped and buffered. The ingestion architecture has already handled device connections and brought related events into shared collection. Consumers can use that data for the processing their applications require.</p>
<p>Destinations can include <a href="https://kafka.apache.org/intro/">Kafka</a> for event distribution, <a href="https://clickhouse.com/docs/get-started/about/intro">ClickHouse</a> for analytics and <a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html">S3</a> for object storage. These services have different downstream roles, but each can begin from the organized event data available at the handoff. The ingestion boundary’s contribution is to prepare that data while accommodating the device traffic that produces it.</p>
<h2 id="make-live-event-traffic-manageable">Make live-event traffic manageable</h2>
<p>The gateway and shared Valkey buffer protect internal processing through cooperating responsibilities. Valkey Router handles device connections and controls ingestion. Batching shares forwarding overhead across small beacons. Deterministic grouping keys bring related events from different gateways into coherent data streams. Shared buffering gives those events room to wait through temporary differences between ingestion and consumption.</p>
<p>Together, these mechanisms let internal services process organized event data while the boundary handles unpredictable device traffic. Their combined effect depends on <a href="https://www.gomomento.com/blog/introducing-valkey-lab-stop-guessing-when-your-cache-hits-its-limit/">sufficient capacity for the traffic</a> and its waiting duration. The architecture makes that relationship explicit, connecting efficient ingestion with the headroom needed to support <a href="https://www.gomomento.com/blog/what-is-real-time-data-processing/">real-time processing</a> during live events.</p>
<p>If you are designing this boundary, bring your expected beacon traffic, payload sizes, grouping requirements and downstream processing needs to <a href="https://www.gomomento.com/contact-us/">Momento</a>. We can help you explore how Valkey Router and shared Valkey collection fit your workload and topology.</p>]]></content:encoded>
    </item>
    <item>
      <title>A practical architecture for remote KV caching with Valkey</title>
      <link>https://www.gomomento.com/use-cases/remote-kv-caching-with-valkey/</link>
      <pubDate>Thu, 01 Oct 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <category><![CDATA[AI]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/use-cases/remote-kv-caching-with-valkey/</guid>
      <description><![CDATA[<p>Reuse KV cache across inference workers with a shared Valkey cache, RDMA/EFA networking, plus NVMe and object storage tiers, all powered by Momento.</p>]]></description>
      <content:encoded><![CDATA[<p>An inference fleet can process the same context many times even when prefix caching is enabled. A worker keeps useful state locally, but the next request lands elsewhere. Long document prefixes get recomputed while their earlier results sit on another machine or have already been evicted.</p>
<p>A shared Valkey tier gives those results a place that every compatible worker can reach. Momento’s RDMA and AWS EFA modules provide a direct path between inference GPUs and KV blocks in Valkey RAM. Its NVMe module adds local disk capacity, and its object storage connector asynchronously flushes data to inexpensive long term storage. These capabilities let you retain reusable state beyond the lifetime and memory budget of an individual worker.</p>
<p>Consider an assistant answering questions about an operations manual. Each prompt begins with the same system instructions and manual, followed by a different question. The objective is to reuse that shared prefix wherever the next question runs.</p>
<h2 id="keep-the-useful-state">Keep the useful state</h2>
<p><a href="https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/">KV cache</a> contains the attention keys and values computed for prior tokens. Prefill processes the input and establishes this state. Decode uses it to generate the response and extends it as new tokens arrive. Keeping those tensors avoids recreating them during generation.</p>
<p><a href="https://www.gomomento.com/blog/kv-caching-pays-off-under-load/">Prefix caching</a> extends reuse across requests. If a later prompt has the same token prefix and compatible model state, the engine can reuse completed cached blocks. It still processes the new question and generates an answer. Matching subject matter is insufficient. Changing the instructions or moving the manual later in the prompt can change the reusable prefix.</p>
<p>Remote reuse also differs from <a href="https://www.gomomento.com/blog/disaggregation-makes-kv-cache-a-system-primitive/">prefill and decode disaggregation</a>. Disaggregation runs the two phases on separate worker pools and requires a state handoff between them. A shared cache can support that design, but it also benefits workers that run both phases together. The manual example uses such workers.</p>
<h2 id="put-valkey-between-inference-and-storage">Put Valkey between inference and storage</h2>
<p>Deploy Valkey on storage nodes separate from the inference workers. GPU memory holds state needed for active inference. Valkey RAM holds reusable blocks shared across the fleet. Local NVMe on the Valkey nodes and an object bucket extend the retained working set.</p>
<p><a href="https://www.gomomento.com/blog/snowflake-moment-for-inference/">Separating storage from compute</a> lets you add cache capacity without buying another inference GPU. It also gives new or replaced workers access to previously computed prefixes. Valkey supplies an open source storage foundation with a <a href="https://valkey.io/topics/modules-intro/">module extension mechanism</a>. Momento’s modules add the transfer and storage capabilities needed for this architecture.</p>
<p>The following reference design stages lower tier blocks through Valkey RAM. Placement, promotion, and coordination are architectural choices for this design.</p>
<p><img src="https://www.gomomento.com/assets/content/use-cases/remote-kv-caching-with-valkey/architecture.png" alt="Remote KV caching architecture: Workers A and B coordinate lookup through the Coordinator and exchange KV blocks with Valkey RAM via the RDMA / EFA module. Shared KV block storage holds hot blocks in Valkey RAM, warm blocks in the NVMe Module, and long-term blocks in the Object Connector."></p>
<p><em>Direct GPU transfers use Valkey RAM. Lower tier blocks return through RAM in this reference design. Object writes run asynchronously, separately from immediate reuse.</em></p>
<p>The runtime integration identifies reusable blocks, coordinates availability, and arranges for tensors to become usable GPU cache state. Lookup and bulk transfer have separate roles. Finding a block does not mean that the GPU is ready to consume it.</p>
<p>Suppose worker A has processed the manual and exported eligible completed blocks. Worker B receives another question and checks its local cache first. For a local miss, the integration looks for compatible prefix blocks in Valkey. A RAM hit uses Momento’s RDMA or EFA module to move the blocks into worker B’s GPU memory. The runtime then processes the unmatched prompt suffix and continues generation.</p>
<p>If no useful blocks are available, worker B computes the prefix. Its new blocks can be admitted to shared storage for later requests. Keep retrieval bounded so a slow lookup does not turn a recoverable miss into a long wait. Once RAM reuse works, the next step is retaining more useful prefixes than RAM alone can hold.</p>
<h2 id="extend-the-working-set-with-nvme-and-object-storage">Extend the working set with NVMe and object storage</h2>
<p>Give each tier a distinct role. Valkey RAM serves shared prefixes used frequently. Momento’s NVMe module provides more capacity on the storage node for blocks worth keeping outside RAM. Object storage provides a larger retention tier for contexts that may return after hours, sessions, or worker replacements.</p>
<p>In this reference design, requested NVMe blocks are staged into Valkey RAM and then transferred to the GPU through the same fast path. The disk is local to Valkey, so inference workers access it through the storage service. They do not each need a disk copy of every manual.</p>
<p>Momento’s object connector flushes data asynchronously. A block already available in RAM can be reused while its object copy is being written. That decouples immediate reuse from the upload, but it does not make every cache write immediately durable. If the only usable copy disappears before the flush completes, the prefix may need recomputation.</p>
<p>For an object hit, the proposed return path reads retained blocks into Valkey RAM before loading the GPU. Recoverable lookup metadata must accompany retained state so the system can identify compatible objects after a restart. Object storage extends the reuse window, while retrieval still has to fit the request’s latency budget. Choose a storage class that supports the access times you need.</p>
<p>Admission and retention policies keep this hierarchy useful. Prefer blocks likely to be reused over a stream of one-time prompts. Keep copies available until transfers using them finish. Coordinate expiry across lookup metadata and lower tiers so deleted or incompatible state cannot appear as a usable hit. Evaluate the tradeoffs between <a href="https://www.gomomento.com/blog/a-roadmap-for-kv-cache-offloading-at-scale/">capacity and coordination</a> when defining your architecture.</p>
<p>Workers still perform attention using state loaded into their GPU cache. The remote tiers preserve reusable blocks between uses. If reading and installing an older prefix costs more than recomputing it, a bounded fallback is the better serving decision. That comparison makes the network part of the cache design.</p>
<h2 id="match-the-network-to-the-workload">Match the network to the workload</h2>
<p>Measure <a href="https://www.gomomento.com/blog/disaggregated-llm-inference-part-3-why-your-networking-stack-may-not-be-ready/">effective block transfer bandwidth</a> under serving load, including lookup, lower tier reads, and GPU installation. Network line rate alone cannot tell you whether a hit will beat prefill. Concurrent inference traffic and cache fills also compete for bandwidth.</p>
<p>RDMA reduces overhead by moving data between registered memory regions without the usual socket data path. With compatible hardware and software, <a href="https://docs.nvidia.com/cuda/gpudirect-rdma/">GPU direct transfer</a> avoids staging the payload through inference host memory. Momento supplies the Valkey side of that direct path.</p>
<p>On AWS, use Momento’s EFA module with compatible GPU and storage endpoints. Confirm the instance pairing and GPU direct support before provisioning. <a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/efa.html">EFA device traffic</a> stays within an Availability Zone and VPC. Place the object bucket in the same Region and access it through normal object APIs.</p>
<h2 id="connect-the-serving-stack">Connect the serving stack</h2>
<p><a href="https://www.gomomento.com/products/cache/">Momento Cache</a> is the leading provider of <a href="https://valkey.io/topics/migration/">Valkey</a>, an open-source fork of Redis built for real-time workloads. With Momento Cache and Valkey, your team can enjoy the same high-performance infrastructure that powers global leaders like Snap, FOX, Coinbase, and Capcom.</p>
<p>Momento provides direct integrations with SGLang, vLLM, and LMCache. Choose the path that fits your serving stack. For the manual example, start with the direct vLLM integration if vLLM already runs your workers. If LMCache already manages offloading, use Momento’s LMCache integration. The direct SGLang integration offers the corresponding entry point for SGLang deployments.</p>
<p>Both workers need the same model revision and adapters, matching tokenized prefixes, and compatible tensor representation and parallel layout. Give a new model or cache format its own namespace. Scope sharing to authorized tenants or trust groups, and protect storage access independently of prefix matching. A content hash alone is not an authorization boundary. These requirements belong in the integration and deployment configuration so application code can stay focused on requests.</p>
<h2 id="how-to-measure-the-impact-of-a-shared-cache">How to measure the impact of a shared cache</h2>
<p>A useful first deployment needs two compatible inference workers, one separate Valkey node with RAM and NVMe, and an object bucket. Obtain the Momento modules and runtime integration for the chosen topology. Start by proving shared RAM reuse before introducing disk and object retrieval.</p>
<p>Send a long manual prefix to worker A, then send a different question with that same prefix to worker B. Keep local prefix caching enabled in your baseline. Confirm that B lacks a local copy and obtains compatible blocks from shared storage. Compare its time to first token with recomputing the prefix under the same serving conditions.</p>
<p>Next, grow the working set beyond the RAM tier to exercise NVMe. Test object retrieval using blocks whose flush has completed and whose faster copies have been removed through supported controls. Restart an inference worker to check that reuse survives worker turnover. Test an unavailable remote tier and verify bounded retrieval followed by recomputation.</p>
<p>Measure tail latency and transfer time as concurrency rises. Track hits and retrieved bytes by tier, GPU memory pressure, and the object flush backlog. Include cache memory, disk, object requests, and network costs when judging the result. A hit rate without its retrieval cost can make an expensive cache look successful.</p>
<h2 id="accelerate-your-kv-caching-with-valkey-and-momento">Accelerate your KV caching with Valkey and Momento</h2>
<p>Valkey is a practical solution when repeated prefixes outlive one worker’s cache and loading their state beats recomputation. Momento’s modules connect that shared RAM tier to inference GPUs and extend it onto disk and object storage. Looking for advice or help to trying it out? Get in touch with <a href="https://www.gomomento.com/contact-us/">Momento</a> and one of our solution engineers will be happy to take a look!</p>]]></content:encoded>
    </item>
    <item>
      <title>50 Million Sorted Sets, Round Three: Redis Still Chasing Valkey</title>
      <link>https://www.gomomento.com/blog/50-million-sorted-sets-round-three-redis-and-valkey-compared/</link>
      <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
      <pubDate>Sun, 27 Sep 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/blog/50-million-sorted-sets-round-three-redis-and-valkey-compared/</guid>
      <description><![CDATA[<p>A 50-million-member sorted set benchmark compares memory use and insert throughput across six Redis releases and five Valkey releases.</p>]]></description>
      <content:encoded><![CDATA[<p>I’ve now run the same sorted set benchmark <a href="https://valkey.io/blog/50-million-zsets/">twice</a> on Valkey: 50 million members pushed into a single ZSET via pipelined ZADD, memory measured at the end. The first post documented Valkey 8.1 cutting memory 22% versus 8.0. The follow-up found another 11% in Valkey 9.1. The obvious next question: what does the same test look like on Redis?</p>
<p>I wanted a reproducible benchmark to assess the progress through the builds. Here is the same benchmark, same tool (<a href="https://github.com/momentohq/sorted-set-benchmark">sorted-set-benchmark</a>, pipelined ZADD NX, members <code>m:{i}</code>, integer scores), same Graviton4 box (r8g.4xlarge), same protocol: ten repetitions per version with a flush in between, official docker images with <code>--save '' --appendonly no</code>, one server at a time, single benchmark client on the same host. Six Redis versions, five Valkey versions.</p>
<h2 id="the-numbers">The numbers</h2>
<p><a href="https://www.gomomento.com/assets/content/blog/50-million-sorted-sets-round-three-redis-and-valkey-compared/benchmark-comparison.svg"><img src="https://www.gomomento.com/assets/content/blog/50-million-sorted-sets-round-three-redis-and-valkey-compared/benchmark-comparison.svg" alt="Combined chart of 11 releases positioned by their version family&#x27;s initial release date, with a minimum gap between nearby releases. Thin orange Redis and green Valkey bars show bytes per member on the left axis, spanning 0 to 150; separate lines show each engine&#x27;s median inserts per second in thousands on the right axis. Angled labels identify the tested engine and patch version. Exact values follow in the table."></a></p>
<div><table><caption>Memory and insert throughput by engine and version</caption><thead><tr><th scope="col">Engine</th><th scope="col">Version</th><th scope="col">used_memory</th><th scope="col">Bytes per member</th><th scope="col">Median inserts/sec</th></tr></thead><tbody><tr><th scope="row">Redis</th><td>7.4.11</td><td>4.83 GB</td><td>~97</td><td>539k/s</td></tr><tr><th scope="row">Redis</th><td>8.0.6</td><td>4.83 GB</td><td>~97</td><td>544k/s</td></tr><tr><th scope="row">Redis</th><td>8.2.10</td><td>4.83 GB</td><td>~97</td><td>551k/s</td></tr><tr><th scope="row">Redis</th><td>8.4.7</td><td>4.83 GB</td><td>~97</td><td>536k/s</td></tr><tr><th scope="row">Redis</th><td>8.6.7</td><td><strong>3.73 GB</strong></td><td><strong>~75</strong></td><td>613k/s</td></tr><tr><th scope="row">Redis</th><td>8.10.2</td><td><strong>3.73 GB</strong></td><td><strong>~75</strong></td><td>632k/s</td></tr><tr><th scope="row">Valkey</th><td>7.2.14</td><td>4.83 GB</td><td>~97</td><td>516k/s</td></tr><tr><th scope="row">Valkey</th><td>8.0.11</td><td>4.83 GB</td><td>~97</td><td>533k/s</td></tr><tr><th scope="row">Valkey</th><td>8.1.10</td><td>3.77 GB</td><td>~75</td><td>590k/s</td></tr><tr><th scope="row">Valkey</th><td>9.0.6</td><td>3.77 GB</td><td>~75</td><td>571k/s</td></tr><tr><th scope="row">Valkey</th><td>9.1.2</td><td><strong>3.34 GB</strong></td><td><strong>~67</strong></td><td><strong>647k/s</strong></td></tr></tbody></table></div>
<p>Memory figures are the server’s <code>used_memory</code> (decimal GB, as the tool reports it). Every version’s ten repetitions landed on the same value at the tool’s reporting precision, roughly 10 MB granularity, with <code>mem_fragmentation_ratio</code> at 1.01x throughout, so RSS tracks these numbers closely. Throughput spreads across the ten runs were tight; the two headline configurations don’t overlap (Valkey 9.1: 643 to 652k/s; Redis 8.10: 628 to 639k/s).</p>
<p>Three things stand out. First, every version of both engines that predates this optimization work sits at the same 4.83 GB for this workload, which is what you’d expect given the shared codebase history of the ZSET implementation. Second, Redis ships a comparable reduction: memory drops from 4.83 GB to 3.73 GB, 23%, exactly at the 8.6 boundary (8.4.7 has the old layout, 8.6.7 the new one), with median inserts rising from 536k/s to 613k/s. Third, Valkey 9.1.2 remains the smallest and fastest configuration I tested: about 10% less memory than Redis 8.10.2 and a small but consistent throughput edge.</p>
<h2 id="the-same-idea-traveling">The same idea, traveling</h2>
<p>I found it really interesting that Redis’s newest number lands within one percent of Valkey 8.1’s: 3.73 GB versus 3.77, about 75 bytes per member either way. Redis in late 2026 is, byte for byte, roughly Valkey of April 2025. That’s not a coincidence, and you can trace exactly why through public pull requests and release notes.</p>
<p><a href="https://github.com/valkey-io/valkey/pull/1427">Valkey PR #1427</a> (merged January 2025, shipped in Valkey 8.1 that April) restructured the sorted set so the hash lookup structure points directly at the skiplist node, instead of keeping its own copy of the mapping. In this benchmark that’s the 4.83 to 3.77 GB step. <a href="https://github.com/valkey-io/valkey/pull/2508">Valkey PR #2508</a> (merged November 2025, shipped in 9.1) then embedded the member string inside the skiplist node itself, removing a pointer and a separate allocation per element. That’s the 3.77 to 3.34 GB step.</p>
<p><a href="https://github.com/redis/redis/pull/14701">Redis PR #14701</a> (merged January 2026, a year and ten days after Valkey’s) copied this work to Redis, and the PR description says where it came from, verbatim: “This is based on: valkey-io/valkey/pull/1427.” Credit to the engineer for saying so plainly; not everyone does. The adaptation kept Redis’s dict (in <code>no_value</code> mode) instead of adopting Valkey’s new hashtable, grafted on the node embedding, and shipped in Redis 8.6. My measurements agree with their release notes: the drop lands exactly at 8.6.</p>
<p>So why is Redis still 10% behind? Because copying is a trailing indicator. Redis took the design but not the hashtable it was built on, and by the time the port shipped, Valkey had already taken the next step. That’s the whole gap: the piece of the design they didn’t adopt, plus the release they haven’t caught up to yet. I mean this without a drop of snark toward the engineer who did the work: Valkey is BSD licensed precisely so anyone, including Redis, can build on it, and Redis users are genuinely better off for it. If the ideas flowed the other way, I’d report that too. But I benchmark what shipped, and what shipped says one project is setting the pace on this data structure and the other is matching it, about a year behind.</p>
<h2 id="the-operators-takeaway">The operator’s takeaway</h2>
<p>Same caveats as always: one large skiplist-encoded ZSET, short members, integer scores, allocator-level accounting, a single pipelined client. Your data shape will land somewhere else, so run the <a href="https://github.com/momentohq/sorted-set-benchmark">tool</a> against your own workload.</p>
<p>If you run Valkey, this is simple: 9.1 posted the lowest memory and the highest insert rate of everything on this table, and the upgrade is the whole optimization. I’ll keep running this benchmark as both projects ship.</p>
<p>If you run Redis and are worried that Valkey is more efficient, you have two options. You can wait for Redis to eventually copy the latest Valkey PR, or you can switch to Valkey. What do you plan to do?</p>
<h2 id="momento-cache-the-fastest-valkey-in-the-cloud">Momento Cache: the fastest Valkey in the cloud</h2>
<p>You always get the latest Valkey through Momento’s managed services, without managing the infrastructure yourself. Momento smoothly handles upgrades, failover, optimization, and more with the same technology that powers some of the world’s largest caches. Join companies like Snap, Coinbase, Paramount, and Capcom: <a href="https://console.gomomento.com/">try Momento Cache now</a>!</p>]]></content:encoded>
    </item>
    <item>
      <title>Your KV cache belongs on the network</title>
      <link>https://www.gomomento.com/blog/your-kv-cache-belongs-on-the-network/</link>
      <dc:creator><![CDATA[Allen Helton]]></dc:creator>
      <pubDate>Mon, 14 Sep 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <category><![CDATA[AI]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/blog/your-kv-cache-belongs-on-the-network/</guid>
      <description><![CDATA[<p>The KV cache is huge, so shipping it over a network should be a non-starter. I benchmarked DRAM, local NVMe, and remote Valkey to find out what the hop really costs, and the answer wasn't what I expected.</p>]]></description>
      <content:encoded><![CDATA[<p>You know how back in the day we used to be messing around at our desk and a teammate asks “<em>What are you doing?”</em> To which our answer was always “<em>I’m just waiting for this to finish compiling.</em>” In 2026, we have a different excuse: “<em>I’m waiting for Claude to respond.</em>”</p>
<p>We say it jokingly, but it’s true. Inference isn’t well-known for being fast. Powerful, sure. But not fast.</p>
<p>We’ve done a lot of experimentation to make it faster. Parallel fetch, zero-copy reads, batched metadata lookups, the whole retrieval path. But there was one particularly intriguing experiment that has a lot of upside, but some huge trade-offs. What if we stored the KV cache remotely instead of locally? There’s just one problem. The KV cache is huge. Transporting it over a network is a non-starter. Right?</p>
<p>We know how it works. GPU memory is fastest, then host DRAM, then the NVMe bolted to the same box, then, if you really have to, go across the network. You go remote when you need to share, but you pay for it in latency.</p>
<p>I’m not onto a novel idea though. Our industry is all over it right now. NVIDIA has an <a href="https://docs.nvidia.com/dynamo/dev/knowledge-base/modular-components/backends/v-llm/kv-cache-offloading">offloading tier in Dynamo</a>. Mooncake is <a href="https://kvcache-ai.github.io/Mooncake/">building a store</a>. Xinnor published a piece in June arguing that Lustre with a client-side cache <a href="https://xinnor.io/blog/kv-cache-storage-is-the-new-ai-inference-bottleneck-how-lustre-ro-pcc-and-xiraid-make-shared-storage-competitive-with-local-nvme-for-long-context-llm-serving/">can hold its own against local NVMe</a>. <a href="https://www.linkedin.com/in/kshams/">Khawaja Shams</a> made the case that once you <a href="https://www.gomomento.com/blog/disaggregation-makes-kv-cache-a-system-primitive/">disaggregate prefill and decode</a>, the cache has to travel between machines whether you like it or not.</p>
<p>Clearly there’s merit to it. But I didn’t see much about the cost of the tradeoff. How much of a latency hit do you really take storing your KV cache in something like Valkey, and is skipping prefill still worth it after you’ve paid it?</p>
<p>So I took it upon myself to run the numbers.</p>
<h2 id="the-setup">The setup</h2>
<p>The benchmark uses three machines: two identical GPU nodes running vLLM, and a Valkey box sitting one 10 Gbps hop away from both of them.</p>
<p><img src="https://www.gomomento.com/assets/content/blog/your-kv-cache-belongs-on-the-network/setup.webp" alt="Benchmark setup: gpu-a computes the KV with DRAM and NVMe tiers inside the box, Valkey holds the shared cache one 10 Gbps hop away, and gpu-b is an identical node whose own tiers start empty"></p>
<p><em>gpu-a</em> does all the computing for prefill. The DRAM and NVMe tiers live inside it, and Valkey sits across the network. <em>gpu-b</em> is used to determine what happens when a cached request lands on a machine that did not perform the prefill.</p>
<p>The workload is thirty <a href="https://github.com/momentohq/kvcache-corpus">realistically generated legal and medical documents</a>, all trimmed to exactly 10,000 tokens. Every document goes to vLLM twice. The first pass does the prefill and stores the KV cache. The second pass sends the same prompt, and I time how long the first token takes to come back (TTFT). The only thing that changes between runs is where the KV cache lives.</p>
<p>Everything runs <a href="https://github.com/vllm-project/vllm/releases/tag/v0.28.0">vLLM 0.28.0</a> with <a href="https://github.com/LMCache/LMCache/releases/tag/v0.5.4">LMCache 0.5.4</a>, <a href="https://pypi.org/project/valkey-glide-sync/">valkey-glide-sync</a> 2.5.1, and Valkey 9.1.1, serving <a href="https://huggingface.co/Qwen/Qwen2.5-7B-Instruct-AWQ">Qwen2.5-7B-Instruct-AWQ</a> on <a href="https://www.nvidia.com/en-us/data-center/l4/">NVIDIA L4s</a>.</p>
<p>To get accurate benchmark results I had to turn off vLLM prefix caching so the GPU wouldn’t serve every repeat request. I also had to clear the disk tier and drop the page cache before every single run, because Linux keeps recently written disk data in memory. If I didn’t do that, the “NVMe” reads would come out of RAM and I wouldn’t have actually tested the disk speed.</p>
<p>The goal was to beat cold TTFT, which falls between 3.1 and 3.3 seconds. I ran the entire benchmark twice on separate instances a day apart, and got reproducible numbers within 2%. So I’m feeling good about what I found. Let’s take a peek and see how the network did.</p>
<h2 id="keeping-everything-on-one-machine">Keeping everything on one machine</h2>
<p>For this first test, <em>gpu-a</em> does all the work. It computes the KV, stores it in whichever tier is being tested, and reads it back on the second pass. Even in the Valkey run it’s <em>gpu-a</em> writing and reading its own cache, just over a wire instead of a local bus.</p>





























<table><thead><tr><th>Tier</th><th align="right">Cached TTFT</th><th align="right">Speedup</th><th>Reuse</th></tr></thead><tbody><tr><td>DRAM (8 GB)</td><td align="right">3345 ms</td><td align="right">0.95x</td><td>0/30</td></tr><tr><td>Local NVMe</td><td align="right">1521 ms</td><td align="right">2.18x</td><td>30/30</td></tr><tr><td>Valkey (remote)</td><td align="right"><strong>557 ms</strong></td><td align="right"><strong>5.98x</strong></td><td>30/30</td></tr></tbody></table>
<p>You’re reading that right. The network was the fastest. And by quite a large margin, at that. It beat the local NVMe by 2.7x on the same machine. The disk is physically inside the box with no switch in the way, and the cache was still quicker to reach from across the network.</p>
<p>(Ignore the DRAM row for now. Something interesting happened there, and we’ll talk about it later.)</p>
<p>At first I thought I broke something. Maybe I accidentally put the “disk” tier on the EBS root volume instead of the instance store? But <code>lsblk</code> showed the tier on <code>nvme0n1</code> (the 250 GB instance store), and <code>df</code> showed 16 GiB written to it during the run, which is about right for 300,000 tokens of KV.</p>
<p>After doing some math, it started making sense.</p>
<p>Each document’s cached KV is 39 chunks of 256 tokens at 57,344 bytes per token, roughly 572 MB (<em>39 x 256 x 57,344</em>). While I was sanity-checking the harness, I ran the DRAM tier with two documents (so everything fit in memory), and hits came back in about 82 ms. A DRAM hit barely transfers anything, which means 82 ms is basically the overhead cost. So any time beyond 82 ms is what I would consider to be transfer time.</p>
<p>That means Valkey’s 557 ms is really about 473 ms of transfer. 572 MB in 473 ms is 1.2 GB/s, which is a 10 Gbps NIC running flat out. Compare that to the disk, which moved 572 MB in ~1,438 ms. That’s 400 MB/s, which is odd, because NVMe should be able to read north of 2 GB/s.</p>
<p>So the remote tier is bound by the network, and the local tier is bound by its own software. Clearly a lot of work went into <a href="https://docs.lmcache.ai/kv_cache/storage_backends/valkey.html">LMCache’s Valkey connector</a> this year. It pulls with eight parallel workers straight into caller-owned buffers. The local disk path hasn’t gotten the same attention yet, and 400 MB/s from an NVMe says there’s plenty of opportunity. I imagine it won’t stay this way for long.</p>
<p>Anyway, from these results we see that putting the KV cache on the network didn’t have the latency penalty I thought it would. It ran as fast as physics allows 🔥. Even if the disk path gets some love next month and starts doing 1.5 GB/s, using the network is still a solid choice. Plus, Valkey has room to go even faster with a bigger NIC.</p>
<p>There is one thing I can’t explain. That 2.18x speedup for the disk is a median, and the per-document numbers show something weird. The first document requested in the cached pass comes back around 11x faster. The second comes back around 3.5x. The other 28 are all right at 2.1x. Oddly enough, both runs were basically the same:</p>

























<table><thead><tr><th>Position in the cached pass</th><th align="right">Run 1</th><th align="right">Run 2</th></tr></thead><tbody><tr><td>1st document</td><td align="right">11.42x</td><td align="right">10.84x</td></tr><tr><td>2nd document</td><td align="right">3.65x</td><td align="right">3.28x</td></tr><tr><td>the other 28</td><td align="right">~2.1x</td><td align="right">~2.1x</td></tr></tbody></table>
<p>Something is warm for the first request or two, then stops. I don’t know what it is, and I’m not going to make up a reason that sounds good. Whatever it’s doing, it makes the disk look better than the median I’m quoting, so it isn’t driving my conclusion.</p>
<h3 id="lets-talk-about-dram">Let’s talk about DRAM</h3>
<p>In my results table above, the DRAM row says it had a 0.95x “speedup”. I know what you’re thinking: I misconfigured it. I didn’t. Remember that two-document sanity check where hits came back in 82 ms? That was this tier with the same config. The TTFT was 3122.8 ms cold and 82.5 ms cached, which is a 37.93x speedup 🤯.</p>
<p>DRAM isn’t slow. DRAM is the fastest thing in the entire stack by a mile. It just… runs out.</p>
<p>Thirty documents is 300,000 tokens of KV. At 57,344 bytes per token, that’s 16 GiB. I gave the tier 8 GB, and the entire host only has 16. The box simply didn’t have the memory, and it evicted everything before I could even run the second pass.</p>
<p>Not only that, but every tier that served nothing came back 3% to 5% <em>slower</em> than cold, on at least 29 of the 30 documents. A lookup that misses is still a lookup, and a store that’s evicting still costs a store. In this instance, the undersized cache was basically a tax.</p>
<p>And thirty documents is nothing when compared to a real production corpus. If your working set is two documents, use DRAM and stop reading this. If it’s thirty or more, skip the DRAM for something with more capacity.</p>
<h2 id="expanding-to-two-machines">Expanding to two machines</h2>
<p>Everything we’ve covered so far has been a single machine writing and then reading its own cache. But the big question is what happens when a request is routed somewhere else, because it will in production. So I ran the cold pass on gpu-a and the cached pass on gpu-b (the one that didn’t do the prefill).</p>





























<table><thead><tr><th>Tier</th><th align="right">One node</th><th align="right">Second node</th><th>Reuse on second node</th></tr></thead><tbody><tr><td>DRAM</td><td align="right">0.95x</td><td align="right">0.96x</td><td>0/30</td></tr><tr><td>Local NVMe</td><td align="right">2.18x</td><td align="right">0.95x</td><td><strong>0/30</strong></td></tr><tr><td>Valkey</td><td align="right">5.98x</td><td align="right"><strong>5.97x</strong></td><td><strong>30/30</strong></td></tr></tbody></table>
<p>Unsurprisingly, NVMe went down to zero reuse and 3.2 second prefill on gpu-b because the KV didn’t exist there and there was nothing to access. It makes NVMe seem a lot less exciting once you start using it in a distributed environment.</p>
<p>That said, Valkey went from 5.98x to 5.97x. Reading from a machine that did not compute the KV cost nothing measurable. The worst document slipped from 5.55x to 3.92x, and the medians are effectively the same. The cache doesn’t belong to either node. And that was the whole point.</p>
<h2 id="my-takeaways">My takeaways</h2>
<p>If your working set fits in host DRAM, use that. A cache hit came back 37x faster than recomputing the prefill. Just be honest with yourself about capacity. Measure your corpus in tokens and multiply by your model’s per-token KV. Don’t eyeball it.</p>
<p>Once it stops fitting, local disk gets you 2.18x on the machine that wrote it, but nothing anywhere else. On a single node that’s a considerable boost. But across a fleet, every entry is isolated to the box that produced it, and your hit rate becomes the odds of getting routed back to the same machine.</p>
<p>The network tier isn’t free on the way in. Cold TTFT ran about 110 to 150 ms slower with Valkey than with DRAM, which is a 3% to 4% cost for storing over the wire instead of in memory. So the tradeoff is that you pay 110-150ms per cache fill in exchange for about 2.7 seconds back on every hit after it.</p>
<p>The most surprising result was that the shared tier was already the faster one here. On this hardware, with this model, at this corpus size, putting the cache on another machine made it quicker. I went into this assuming I’d be arguing that sharing is worth a latency cost. I am happy to have been wrong on that one. Skipping prefill is worth what it costs (and then some).</p>
<p>Your KV cache belongs on the network. Not because it has to be shared (though it does). Because that’s where it was fastest.</p>
<h2 id="try-it-yourself">Try it yourself</h2>
<p>The harness, the 30-document corpus, the tier configs and the infrastructure scripts are all in <a href="https://github.com/ksmotiv8/valkey-kvcache-bench">valkey-kvcache-bench</a>. One command provisions all three hosts, runs the matrix, pulls the results down and terminates everything:</p>
<pre><code class="language-powershell">cd infra
.\Run-TierBench.ps1 -AwsProfile &#x3C;profile> -Teardown
</code></pre>
<p>It costs about $2.14 an hour and the full benchmark run takes under an hour.</p>
<p>One warning if you try this yourself - the first time I ran the two-machine test, every tier came back 0/30 (including Valkey). It turns out LMCache’s default cache key hash is Python’s <code>hash()</code>, which gets a different random salt in every process. So <em>gpu-a</em> and <em>gpu-b</em> were generating different keys for identical tokens, which made cross-node reuse impossible. LMCache even warns you at startup with a <em>Using builtin hash without PYTHONHASHSEED set</em> message, and I had completely missed it. But you can fix it by setting this config value:</p>
<p>​<code>yaml pre_caching_hash_algorithm: sha256_cbor_64bit ​</code></p>
<p>If your shared KV cache ever shows a zero hit rate across nodes, it’s probably that.</p>
<p>Anyway, try it out and if you see something different, please feel free to reach out so we can compare.</p>
<p>Happy coding!</p>]]></content:encoded>
    </item>
    <item>
      <title>Momento Cache Cluster and Flex enter limited preview</title>
      <link>https://www.gomomento.com/blog/momento-cache-cluster-and-flex-limited-preview/</link>
      <dc:creator><![CDATA[Dylan Abraham]]></dc:creator>
      <pubDate>Wed, 02 Sep 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <category><![CDATA[News]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/blog/momento-cache-cluster-and-flex-limited-preview/</guid>
      <description><![CDATA[<p>Momento Cache puts high-performance Valkey at your fingertips. Cluster provides direct control over topology, while Flex automatically optimizes resources.</p>]]></description>
      <content:encoded><![CDATA[<p>We’re opening a limited preview of the new <strong>Cluster</strong> and <strong>Flex</strong> configurations for <strong>Momento Cache</strong>. This release conveniently packages the technology and operating experience that power some of the largest Valkey clusters in the world.</p>
<p>Momento Cache is built for fast-moving teams who want optimal Valkey performance without the hassle of babysitting infrastructure. Cluster capacity and Flex capacity bring single-tenant, fully-managed resources to this self-service infrastructure platform.</p>
<p>Cluster provides direct control over instance type, shards, replicas, and availability zones. Flex automatically optimizes resources within specified bounds. Both were designed for workloads that need <strong>stronger isolation</strong>, <strong>more control</strong>, and a <strong>higher performance ceiling</strong> than Momento Cache’s Serverless configuration.</p>
<h2 id="avoiding-the-trap-of-operational-creep">Avoiding the trap of operational creep</h2>
<p>In the AI era, it’s trivial to stand up basic infrastructure at near-zero cost. We’ve all been there: it’s fast, it’s easy, it mostly works. It’s the right solution when you need to ship.</p>
<p>Then, growth hits. Traffic changes shape. Memory fills unevenly. A shard needs to move. A primary fails. Clients stampede. A zero-day patch lands at midnight. Suddenly, <strong>operational creep</strong> has consumed the time and the token budget that you wanted to spend on building.</p>
<p>And the problems only multiply as you add more features, products, and services. Soon your entire team is stuck fighting against infrastructure as latency, cost, and complexity steadily creep up and to the right.</p>
<h2 id="fast-reliable-efficient---pick-all-three">Fast, reliable, efficient - pick all three</h2>
<p>The Momento platform powers critical features for millions of users around the world at companies like Capcom, Coinbase, Paramount, and Snap. Now, Momento Cache puts the full power and flexibility of Valkey at your fingertips, packaging up hard-won production lessons into a streamlined service.</p>
<p>The new Cluster and Flex configurations help you to tailor a Valkey deployment to fit your specific needs. Momento seamlessly operates the lifecycle behind that system: provisioning, health, failover, rolling topology changes, version upgrades, and security patches.</p>
<p>Whether you’re pushing one thousand or one million requests per second, Momento delivers unrivaled performance and resource utilization. As you grow, your infrastructure grows alongside you. For companies with a mature platform org, Momento Cache can even be deployed in a BYOC configuration, as part of your internal developer platform.</p>
<h2 id="transparent-pricing">Transparent pricing</h2>
<p>Momento Cache pricing is designed to be simple, predictable, and cost-effective. Pricing in us-east-1:</p>
<ul>
<li><strong>Flex</strong> starts at $13 per GB-month of physical Valkey storage</li>
<li><strong>Cluster</strong> applies a 25% surcharge to the list price for every deployed instance</li>
<li><strong>Data transfer</strong> includes 200 GB of ingress + egress each month, then $0.05/GB</li>
</ul>
<p>Momento Cache supports rapid autoscaling within a specified capacity range, making it easy to reduce the cost of idle resources.</p>
<h2 id="designing-for-scale">Designing for scale</h2>
<p>Momento Cache employs a two-tiered architecture that significantly improves reliability and efficiency at scale. Each valkey cluster sits behind a gateway that handles the hard traffic problems like hot keys and connection storms before they hit the data layer.</p>
<p>The gateway exposes a RESP endpoint, so Momento Cache is compatible with all standard Valkey and Redis clients. It masks the underlying cluster, presenting a single stable endpoint across any topology changes.</p>
<pre><code class="language-mermaid">block
  app("Redis or Valkey\nclient"):3
  space
  gateway("gateway"):3
  space
  pool("Valkey cluster"):3

  app --> gateway
  gateway --> pool
</code></pre>
<p>The gateway is optimized to quickly process TLS, auth, rate limits, and other traffic management concerns at high concurrency and high throughput. It multiplexes client traffic across a pool of warm connections to the Valkey cluster. This reduces connection latency, and enables advanced capabilities like request coalescing.</p>
<p>The result is an efficient system with deliberate separation of responsibilities. The gateway absorbs traffic concerns, while Valkey nodes focus their resources on handling data.</p>
<h2 id="get-started-with-momento-cache">Get started with Momento Cache</h2>
<p>Try out Momento Cache in a few short steps with the <a href="https://github.com/momentohq/momento-cli">Momento CLI</a>. While the service is still in preview, you’ll also need to request access via the web console.</p>
<p>First, create an API key in the <a href="https://console.gomomento.com/">Momento console</a>, copy the endpoint for your region, and configure the default CLI profile:</p>
<pre><code class="language-sh"># paste the api key and endpoint when prompted
momento configure
</code></pre>
<p>Then, create a Capacity Pool. Be sure to provide valid zone IDs for your region:</p>
<pre><code class="language-sh">momento preview pool create \
  --name example-pool \
  --capacity-gib 32..128 \
  --replicas-per-shard 1..2 \
  --zones use1-az1,use1-az2

momento preview pool describe --name example-pool
</code></pre>
<p>Once the Capacity Pool status is <code>active</code>, you can create a Database:</p>
<pre><code class="language-sh">momento preview database create \
  --name example-db \
  --pool-name example-pool
</code></pre>
<p>Back in the console, open the pool’s <strong>Databases</strong> tab and copy the regional RESP endpoint. Connect any Valkey or Redis client in standalone mode to this endpoint over TLS on port <code>6379</code>. The Database name is the username, and your Momento API key or token is the password. Here, we’ll demonstrate with the official <a href="https://valkey.io/topics/cli/">valkey cli</a>:</p>
<pre><code class="language-sh">valkey-cli -h &#x3C;resp-endpoint> -p 6379 --tls \
  --user example-db --pass &#x3C;momento-api-key>

> SET example-key "ready"
OK
> GET example-key
"ready"
</code></pre>
<p>Congratulations! You now have a high-performance Valkey cluster ready to go.</p>
<p>Next, check out the <a href="https://docs.momentohq.com/product/cache/">docs</a> to learn more about Momento Cache’s capabilities. Or, if you want to push your cache to the limit, load up <a href="https://github.com/cachecannon/cachecannon">cachecannon</a> on a <code>c7g.xlarge</code> instance in the same region and zone as your Database!</p>
<h2 id="up-next">Up next</h2>
<p>This launch begins the next chapter for Momento Cache. Stay tuned as we port more features from our enterprise services into Momento Cache, including VPC peering and S3 integration!</p>
<p>We’re looking for feedback from teams operating demanding Valkey workloads as we refine the product. If you’re building a fast-growing product, operating a large Valkey or Redis cluster, designing an internal caching platform, or helping teams adopt Valkey, we would love to hear what you need next and where we can help out.</p>
<p>We’re grateful to the engineers and partners who turned years of demanding operating experience into a service anyone can start using today. Special thanks to <strong>Dylan Abraham</strong> and <strong>Jason LaPier</strong> for leading the development effort!</p>]]></content:encoded>
    </item>
    <item>
      <title>Stop counting indexes</title>
      <link>https://www.gomomento.com/blog/stop-counting-indexes/</link>
      <dc:creator><![CDATA[Allen Helton]]></dc:creator>
      <pubDate>Thu, 27 Aug 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/blog/stop-counting-indexes/</guid>
      <description><![CDATA[<p>Fifty extra indexes don't slow down your writes, but one vector field slows them down 11x. I wanted to figure out why valkey-search behaves this way, and it wasn't what I expected.</p>]]></description>
      <content:encoded><![CDATA[<p>Free text search is a beast. Sometimes a user will type in a description of what they’re looking for, like “wireless earbuds under $100,” and other times they’ll copy/paste a SKU for an item they’re replacing. Both are completely valid use cases, but the former requires a vector search and the latter relies on lexical matching. I’ve been spending a lot of time in <a href="https://valkey.io/topics/search/">valkey-search</a> trying to handle both. As you can probably imagine, one index doesn’t cut it, but two indexes on the same keys can.</p>
<p>But doubling the number of indexes I was using made me nervous. I didn’t know what impact that would have on performance. I’m using Valkey because I needed ultra-low latency, I didn’t want to shoot myself in the foot trying to be clever.</p>
<p>And yes, going from one to two indexes is probably not a big deal. But indexes have a way of piling up. You create one to filter products by category, then add price to it a month later, then the search team wants a vector field, then the payment team indexes their keys, then fulfillment indexes theirs. Before you know it, there are a dozen <code>FT.CREATE</code> statements in the repo with no single owner.</p>
<p>My instinct tells me that performance scales inversely with the count. Every index in valkey-search subscribes to a key prefix. When you write a matching key, that index gets an entry in a mutation queue, and your client stays blocked until every entry is processed. That’s what gives you read-after-write consistency. If you have a dozen indexes, you also have a dozen entries, and a dozen times the wait. Right?</p>
<p>So I built a benchmark to see how much additional indexes cost in valkey-search. And I discovered I was very wrong.</p>
<h2 id="the-setup">The setup</h2>
<p>The server was an <code>r7i.4xlarge</code> with 8 physical cores, running Valkey 9.1.1 and valkey-search 1.2.1. valkey-search sized its own writer pool to 8 threads. I turned persistence off, because I didn’t need a background save forking mid-run throwing a wrench into my latency numbers. Load came from a separate <code>c7i.8xlarge</code> in the same placement group, over 384 connections.</p>
<p>Every write is the same 4.6KB product hash that includes a category, a price, a SKU, a title, a description, and a 1024-dimension embedding. The load generator sends on a schedule and times each request from when it was supposed to go out, not when it managed to (more on this later). Every point runs for 25 seconds, three times, and I capture the median.</p>
<h2 id="do-idle-indexes-cost-anything">Do idle indexes cost anything?</h2>
<p>As I mentioned earlier, indexes subscribe to a key prefix. If a key is upserted that doesn’t match on an index, does that slow things down? In other words, my test here determines if the existence of indexes slows down the performance of others.</p>
<p>I ran the same write workload against <code>product:</code> keys in two setups: one index on <code>product:</code>, and another one with the same index plus fifty more indexes on unrelated prefixes.</p>




















<table><thead><tr><th>setup</th><th>sustained writes/sec</th><th>p99 @ 2,000/s</th></tr></thead><tbody><tr><td>1 index on <code>product:</code></td><td>32,000</td><td>2.18ms</td></tr><tr><td>same index + 50 on other prefixes</td><td>32,000</td><td>2.17ms</td></tr></tbody></table>
<p>Nice! I couldn’t measure a difference between them at any rate I tried.</p>
<p>This is because the prefix subscriptions live in a <a href="https://en.wikipedia.org/wiki/Trie">trie</a>. When you write <code>product:88213</code>, valkey-search walks the trie and only notifies the indexes whose prefix matches. An index subscribed to <code>orders:</code> doesn’t hear about it. <code>FT.INFO idx:other0</code> reports a field called <code>mutation_queue_size</code>, and across the entire run it never left zero, which validates the promise of the trie architecture.</p>
<p>So that hodgepodge of indexes in your repo isn’t what’s slowing your writes down. Only indexes whose prefix matches the keys you write are in the path at all.</p>
<h2 id="the-impact-of-matching-indexes">The impact of matching indexes</h2>
<p>So what’s the performance impact of having multiple matching indexes? I ran the same workload again, this time adding multiple identical indexes directly on the <code>product:</code> prefix.</p>






























<table><thead><tr><th>indexes on <code>product:</code></th><th>writes/sec</th><th>indexing jobs/sec</th></tr></thead><tbody><tr><td>0</td><td>80,000</td><td>0</td></tr><tr><td>1</td><td>33,636</td><td>33,636</td></tr><tr><td>2</td><td>20,000</td><td>40,000</td></tr><tr><td>4</td><td>16,818</td><td>67,272</td></tr></tbody></table>
<p>Every matching index gets an indexing job, and the write isn’t done until all of them are. To figure this out, I took Valkey’s <code>search_ingest_hash_keys</code> counter and divided by the writes that completed in the same window. I took this number and multiplied it by the write rate to calculate the indexing jobs/sec column.</p>
<p>The expensive part, naturally, is switching indexing on in the first place. Going from zero indexes to one cost me 58% of my write rate. The second index cost another 41%. Doubling from two to four only cost 16%. Each additional index cost less than the previous one.</p>
<p>Adding indexes pulls more total indexing work per second out of the same server, 33,636 jobs a second at one index and 67,272 at four. So one or two indexes clearly weren’t saturating the writer pool, because it eventually went on to do twice the work.</p>
<p>Indexing jobs fan out across a pool sized to your physical core count, and the blocked-client handles collapse into a single block on your connection. The write waits for the slowest single job, which is why four indexes don’t cost four times what one does.</p>
<p>Unfortunately I don’t know what the limiting factor is. My hunch is that it’s the per-index bookkeeping that happens on the main thread before a job reaches the pool, but I didn’t measure that, so 🤷.</p>
<p><img src="https://www.gomomento.com/assets/content/blog/stop-counting-indexes/fanout-p99.webp" alt="p99 write latency against offered write rate, for zero, one, two and four indexes on the product: prefix. All four curves are within about a millisecond of each other below 10,000 writes per second, then separate and turn sharply upward, each at a different rate."></p>
<p><em>NOTE - Adding matching indexes won’t appear to add latency until it’s too late. At 2,000 writes/sec, one index and four indexes were within a millisecond of each other. You won’t catch it watching p99 on a healthy system, because what you’re spending is headroom. You find out it’s gone when you need it.</em></p>
<h2 id="vector-fields-hit-different">Vector fields hit different</h2>
<p>The benchmarks above use TAG and NUMERIC fields, just the normal filter field types. So I went back to the first benchmark, added a single 1024-dimension HNSW vector field to the one-index setup, and re-ran it to find some staggering results.</p>
<p>32,000 writes per second became 2,828. 🤯</p>
<p>That’s an 11x difference in throughput because of a single field. My napkin math says that’s 2-3 milliseconds of CPU per vector insert spread across eight writer threads. To make matters worse, the performance cliffs. I increased the rate on my benchmark by ~40%, and the wheels fell off.</p>




















<table><thead><tr><th>offered writes/sec</th><th>p99</th><th>writer queue depth</th></tr></thead><tbody><tr><td>2,828</td><td>5.65ms</td><td>4</td></tr><tr><td>4,000</td><td>2,046ms</td><td>376</td></tr></tbody></table>
<p>The performance hit goes from five milliseconds to two seconds. The queue went from basically empty to 376 entries deep, and it stayed there for every rate I tried above that. You’re either under the line and fine, or over it and everything is late. If you’re capacity planning, be sure to check whether an index on your hot prefix has a vector field.</p>
<h2 id="do-vector-fields-affect-other-matching-indexes">Do vector fields affect other matching indexes?</h2>
<p>Back to the search feature I was building that required two indexes. A single index can’t do both jobs because of <a href="https://valkey.io/topics/search-data-formats/#stop-word-removal">stop words</a>. The text pipeline is configured for the entire index, and it splits words on punctuation before dropping anything in the stop word list. So <code>IT-500</code> becomes <code>it</code> and <code>500</code>, <code>it</code> is a stop word, so just <code>500</code> is added to the index. Turning on <code>NOSTOPWORDS</code> means your description field will index every <em>the</em>, <em>is</em>, <em>and</em>, and <em>it</em> (plus a lot more) in the catalog. So we need two indexes to have it turned on for one and off for the other.</p>
<p>But that made me wonder what the second index costs when the first one has a vector field. We saw how much of a hit it made to throughput in our earlier benchmarks.</p>




















<table><thead><tr><th>configuration</th><th>sustained writes/sec</th><th>p99</th></tr></thead><tbody><tr><td>semantic (TAG + NUMERIC + HNSW)</td><td>2,828</td><td>5.29ms</td></tr><tr><td>semantic + exact (TEXT, NOSTOPWORDS)</td><td>2,828</td><td>5.44ms</td></tr></tbody></table>
<p>About a three percent difference. The exact match index processed 141,408 text fields during the benchmark run. Both mutations hit the pool together, the write waits for the slower one (the vector). So if you’re already paying the HNSW tax, the lexical index has essentially no additional latency.</p>
<h2 id="my-takeaways">My takeaways</h2>
<p>So it turns out I was asking the wrong question when I started this experiment. I thought the number of indexes I had was going to slow performance down to a crawl. But it doesn’t. The real question is <em>what fields are inside the indexes that match your keys</em>? Which is a relief, when I think about it.</p>
<p>I also learned a couple of things about benchmarking while I was busy answering the wrong question. 😅</p>
<p>Every sustained writes/sec number in this post could have been bigger. At the top rung of my ladder, the no-index setup completed 128,000 writes a second (but my tables show 80,000). At that run rate, it was making every request wait 1.9 seconds in a queue first. A server running at full utilization drains as fast as it fills, so throughput looks perfect, but at a cost to latency. So the number I used in every table is the highest rate where p99 stayed within reason.</p>
<p>It took me three runs to believe what I saw in that four-row table early in this post. The first run stepped the rate by 1.4x per rung, which put two indexes and four indexes on the same 16,000 rung. Which at first made me think indexes three and four had no performance implications. But in reality, they had both fallen apart somewhere between rungs and the ladder couldn’t show me where. So I re-ran it with 1.19x steps and got a cleaner separation. Then I noticed that ladder started at 14,000, which is already well up the curve, so it never measured what latency looks like when the server is idle. My rule for picking a sustainable rate is relative to that idle number, which meant the second run was grading itself on a curve. The third run started at 2,000 and is the one in the table. A rate ladder can’t resolve a difference smaller than its own step, and it can’t tell you where the knee is if it never saw the flat part before it.</p>
<h3 id="try-it-yourself">Try it yourself</h3>
<p>I have the benchmark, scripts, and results <a href="https://github.com/momentohq/valkey-index-bench">available in GitHub</a>. If you want to check my numbers (or disagree with them!), please do and let me know what you find.</p>
<p>If you want to run the same tests on your own Valkey cluster, you can run these three commands:</p>
<pre><code>FT.INFO &#x3C;index> # shows mutation_queue_size (write backlog) for the specified index
INFO search # shows search_writer_queue_size for the whole pool
CONFIG SET search.info-developer-visible yes # unlocks per-field-type counters like search_ingest_field_vector
</code></pre>
<p>It’s cheap to experiment with your existing clusters because these mutations are reversible. <code>FT.DROPINDEX</code> is instant and your data was never in the index to begin with. Add a shape, measure it, then discard it.</p>
<p>Happy coding!</p>]]></content:encoded>
    </item>
    <item>
      <title>Agent Memory on Valkey</title>
      <link>https://www.gomomento.com/blog/agent-memory-on-valkey/</link>
      <dc:creator><![CDATA[Allen Helton]]></dc:creator>
      <pubDate>Wed, 19 Aug 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <category><![CDATA[AI]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/blog/agent-memory-on-valkey/</guid>
      <description><![CDATA[<p>A filter that looks like a query detail can change how Valkey searches your vectors. By the time you notice, the important decisions are already behind you.</p>]]></description>
      <content:encoded><![CDATA[<p>Agent memory sounds like a cut-and-dried vector search problem. You embed the task, find the nearest memories, and give them back to the model. Done.</p>
<p>Unfortunately it’s not that simple. Memories need to be similar, yes, but they also need to be recent. You don’t want memories from a year ago influencing your agent. And outcome is important too. If one approach worked and another failed, you probably want the successful one steering behavior.</p>
<p>I ran into this while building a small demo that puts <a href="https://github.com/momentohq/valkey-agent-memory-demo">Valkey Search inside an agent’s inference loop</a>. Every task the agent finishes gets written as a memory. It persists the task text, the approach that worked, whether it succeeded, when it happened, and a vector of the task. Next time a similar task shows up, the agent recalls what worked instead of rediscovering it.</p>
<p>I ended up learning how valkey-search handles this the hard way, because my first <code>FT.SEARCH</code> call did not work the way I expected:</p>
<pre><code class="language-bash">FT.SEARCH idx:memory "(@outcome:{success} @created_at:[1782900000 +inf])=>[KNN 3 @vector $vec]" PARAMS 2 vec &#x3C;query-vector> DIALECT 2
</code></pre>
<p>Left of the <code>=></code> is an ordinary boolean filter of a tag and a numeric range over two fields I indexed alongside the embedding. Right of it is the vector search.</p>
<p>I thought Valkey was going to find the nearest vectors and apply the tag and timestamp filters afterward. But it doesn’t work like that (in a good way). The filter criteria are actually inputs to the query planner. Depending on how many memories they match, Valkey chooses a different algorithm to search the vector index. In other words, adding a filter changes how the search runs, which meant I needed to rethink my initial memory schema. Turns out my retrieval policy was also an index-design decision.</p>
<h2 id="nobody-post-filters-anymore">Nobody post-filters anymore</h2>
<p>If you’ve used Pinecone, Qdrant, Weaviate, or Milvus, then this should be familiar. All of them decide at query time whether to walk the graph with your filter applied or abandon the graph and brute-force the matching subset instead. Pinecone calls it single-stage filtering. Qdrant calls it query planning. Weaviate calls it a flat search cutoff. Milvus doesn’t really call it anything, it just does it 😂. valkey-search is the same idea, and on the <code>FT.SEARCH</code> vector path it doesn’t implement post-filtering at all.</p>
<p>If you’re a <a href="https://github.com/pgvector/pgvector">pgvector</a> user, however, filtering happens after the index scan. With the default <code>hnsw.ef_search</code> of 40, a filter that keeps roughly 10% of those candidates might leave you with only 4 results. <a href="https://github.com/pgvector/pgvector?tab=readme-ov-file#iterative-index-scans">Iterative index scans</a> were added in 0.8.0 to make that a little better, and they’re off by default.</p>
<p>valkey-search calls it a <a href="https://github.com/valkey-io/valkey-search/blob/main/src/query/planner.cc">query planner</a>. It estimates how many keys your filter matches and picks one of two algorithms to perform the search.</p>
<p>If the estimate is small compared to the index, it pre-filters results by walking the qualified key set, computing each distance directly, and keeping a top-k heap. The <a href="https://en.wikipedia.org/wiki/Hierarchical_navigable_small_world">HNSW graph</a> is never traversed. So valkey-search is essentially doing a brute-force scan over a tiny set.</p>
<p>If the estimate is large, it performs an inline filter. Your predicate is handed to hnswlib as an <code>isIdAllowed</code> functor and evaluated during traversal of the base layer. Non-matching nodes are still visited and expanded, they just don’t get added to the result set.</p>
<p>The cutoff point for one algorithm vs the other is 0.001. So if the number of estimated matching keys is at most .1% of the number of vectors in the index, it will go the brute-force route. Otherwise it uses the inline filter.</p>
<p>For reference, the cutoff point for Milvus is around 7%, which makes valkey-search about 70x stricter. You end up on the inline path more often than you’d think.</p>
<h2 id="make-the-filter-disappear">Make the filter disappear</h2>
<p>On my demo index with a few hundred memories, the threshold works out to well under 1 key, so every recall that matches anything takes the inline path. You’d never notice either way at that scale.</p>
<p>But what would happen with the same schema in production with 1,000,000 memories? The threshold is 1,000 keys. The planner considers the selectiveness of the filter. If <code>@outcome:{success}</code> matches 70% of the index and your <code>@created_at</code> window is also broad, you’re squarely in the inline path. And the planner is right to put you there. If you scope recall to a single tenant with 400 memories, you drop under the threshold, and the query becomes an exact scan. Perfect, fast recall because 400 distance computations is nothing.</p>
<p>Let’s make it more difficult. A filter matching 1% of a large index sits 10x above the cutoff point, so it takes the inline path, where HNSW traverses a significant number of candidates for every one it’s allowed to keep. That results in a lot of extra latency you didn’t account for. And you can’t change that by tuning it, because <code>search.prefiltering-threshold-ratio</code> is immutable unless <code>search.debug-mode</code> is on, and it isn’t in the public configurables table at all.</p>
<p>You can check which path you’re actually on, by the way. valkey-search counts both, named <code>search_prefiltering_requests_count</code> and <code>search_inline_filtering_requests_count</code>. They’re module fields, so you need <code>INFO SEARCH</code> and not plain <code>INFO</code>. You can run your recall query a hundred times to see which one moves.</p>
<p>So your best bet is the index itself. Make the filter act like a namespace. Put a hash tag in the index name, prefix the keys to match, and each tenant’s recall hits one shard against a small index where the filter is mostly irrelevant. Valkey enforces it in both directions, too. A tagged index name requires every prefix to carry the same tag, and an untagged one requires that none of them do. Be careful with this though, because it’s not easy to undo if you change your mind since there’s no <code>FT.ALTER</code> here.</p>
<h2 id="forgetting-is-expensive">Forgetting is expensive</h2>
<p>Removing a vector from an HNSW index calls <code>markDelete</code> and returns. The node isn’t removed from the graph, it’s still visited, it still routes other searches through itself, it just fails the deleted check. <code>search.hnsw-allow-replace-deleted</code> would let the next insert reuse that slot, but it’s false by default and more of a non-production flag. In production the space is stranded until you drop the index.</p>
<p>Luckily, updates are a different story. Modifying an indexed vector routes to an in-place update under the existing label, so the node keeps its slot and just gets its links rewired. Writing a hash field that isn’t the vector doesn’t mess with the graph, and rewriting the vector with identical bytes short-circuits before it gets there.</p>
<p>So updates are cheap, but deletion is where it gets expensive. That includes <code>DEL</code>, eviction, and expiry.</p>
<p>That’s a little scary, because expiry <em>is</em> the recency policy. valkey-search subscribes to generic, expired, and evicted keyspace notifications, so a TTL’d memory really does leave the vector index when the key goes away. Which is what you want, but it’s also what makes graph nodes stranded.</p>
<p>Which means you have to change how you key your memories. Creating a new key every run feels like a safe and reasonable default, and it’s what my demo does with its <code>memory:&#x3C;id></code> per completed task. But add a TTL to keep things fresh, and every expiring key leaves a node behind. That’s an expensive default at scale. Instead, use a stable key derived from a task fingerprint, and update it in place as the agent learns more about that kind of task. This means the task costs only one node instead of one per attempt.</p>
<p>There are two recency controls here to consider. The <code>@created_at</code> range decides what you’re willing to believe on any given query, and it doesn’t delete anything. The TTL decides what you’re willing to pay to store. Range should be your first lever and the TTL should be the slower fallback.</p>
<h2 id="decide-before-you-have-data">Decide before you have data</h2>
<p>This demo surprised me twice: the filter I wrote as a query detail decides which algorithm runs, and the TTL I (almost) added to keep memories fresh led to stranding the index.</p>
<p>Everything here is discoverable, at least. The planner code is a short function to read and understand. The filtering comes from hnswlib. Even the deleted node behavior is a comment in the source, and the counters are available in <code>INFO SEARCH</code>. You don’t have to take my word for any of it (but you should 😜).</p>
<p>What you can’t read your way out of is when you have to make decisions. How the index is scoped and how a memory is keyed are day-one calls, made before you have a single memory to check them against. Outside of that dev-only flag, a rebuild is the only way to reclaim space that deletion stranded. Get those wrong and you’re reindexing.</p>
<p>The demo <a href="https://github.com/momentohq/valkey-agent-memory-demo">is on GitHub</a> if you want somewhere to start. <code>docker compose up -d</code> gets you Valkey with the search module and an agent that writes its own memories. Run a few tasks through it, then look at <code>INFO SEARCH</code> and find out which path your queries are actually taking.</p>
<p>Happy coding!</p>]]></content:encoded>
    </item>
    <item>
      <title>Consistency compounds: Valkey's journey to 200 Gbps</title>
      <link>https://www.gomomento.com/blog/consistency-compounds-valkeys-journey-to-200-gbps/</link>
      <dc:creator><![CDATA[Khawaja Shams]]></dc:creator>
      <pubDate>Thu, 23 Jul 2026 18:00:00 GMT</pubDate>
      <category><![CDATA[Valkey]]></category>
      <guid isPermaLink="true">https://www.gomomento.com/blog/consistency-compounds-valkeys-journey-to-200-gbps/</guid>
      <description><![CDATA[<p>Across four releases, I/O-threading changes removed a serial copy bottleneck, cut p99 latency, and brought large GETs to line rate on our 200 Gbps test rig.</p>]]></description>
      <content:encoded><![CDATA[<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/read-bandwidth.svg" alt="Valkey GET bandwidth by value size across versions 7.2, 8.1, 9.0, and 9.1, with the 200 Gbps NIC limit marked"></p>
<p>If you came here looking for a post from me on consistency models, I have to disappoint you. Today, I want to talk about the value of consistency in life. In life, effort is additive, but consistency is multiplicative.</p>
<p>Over the last few years, I have had the honor of watching the Valkey project blossom from an idea into an inspirational, community-driven effort with consistent improvements in each release. These improvements range from memory efficiency to availability at scale to substantial performance gains. These small improvements compound.</p>
<p>This compounding effect and the power of a driven community is perhaps best illustrated by the journey of the I/O-threading architecture and its impact on large objects from 1 MB to 64 MB in Valkey. Larger items are becoming increasingly important for inference KV caches, 4K video streaming, and other enticing use cases. At the very least, the impact of large objects should not be overlooked, as they can <a href="https://www.gomomento.com/blog/large-objects-in-valkey-9-0/">impact everyone else’s latency</a>; this post measures how fast the large objects themselves go.</p>
<p>Buckle up, because it’s about to get spicy.</p>
<h2 id="a-brief-history-of-io-threads">A brief history of I/O threads</h2>
<p>Valkey 7.2 inherited an I/O-thread model from a software stack doing its best to stay single-threaded. The main thread and I/O threads worked in coordinated phases, separated by synchronization barriers, and the <a href="https://github.com/valkey-io/valkey/blob/7.2/valkey.conf">7.2 config file</a> is candid about the result: “Usually threading reads doesn’t help much.”</p>
<p>Valkey 8.0 delivered a fundamental rearchitecture of I/O threads, <a href="https://valkey.io/blog/unlock-one-million-rps/">tripling throughput to over a million requests per second</a>. This idea was not new. In August 2023, seven months before the fork, Dan Touitou filed <a href="https://github.com/redis/redis/issues/12489">redis#12489</a>, laying out exactly this design in detail, benchmarks included.</p>
<blockquote>
<p>Redis let #12489 sit. The issue is still open in the tracker today, unassigned and without a milestone.</p>
</blockquote>
<p>Within days of the fork, the Valkey community copied the proposal verbatim into <a href="https://github.com/valkey-io/valkey/issues/22">issue #22</a>, greeted it as “a true gem,” and shipped it. The new architecture enabled continuously running I/O threads connected by queues, so reads, parses, and writes proceed on separate cores while the main thread executes commands. Valkey 8.1 delivered TLS handshake offload to I/O threads in <a href="https://github.com/valkey-io/valkey/pull/1338">#1338</a>.</p>
<p>Redis shipped a <a href="https://redis.io/blog/redis-8-0-m03-is-out-even-more-performance-new-features/">strikingly similar asynchronous I/O-threading model</a> in Redis 8 in May 2025, a year after the fork and 21 months after the design landed in its own tracker. Around the same time, we put <a href="https://www.gomomento.com/blog/valkey-turns-one-how-the-community-fork-left-redis-in-the-dust/">Valkey 8.1 and Redis 8.0 head to head</a> on small objects. Valkey 8.1 outran Redis 8.0 by 37% on writes and 16% on reads.</p>
<p>Valkey 9.0 brought <a href="https://github.com/valkey-io/valkey/pull/2078">reply copy avoidance</a>, changing the performance of large items entirely. Before this change, the main thread copied the entire object into a connection reply buffer before moving to the next command. While a large item is being copied, the entire pipeline stalls. Small objects are not serviced until the copy finishes.</p>
<p>Valkey 9.0 instead passes a reference to the I/O threads and keeps the object alive with a reference count. The I/O worker for that connection hands the object’s memory directly to <code>writev()</code>. This shortens the handoff to the I/O threads and gets the main thread back to handling requests.</p>
<p>Valkey 9.1 then redesigned communication between the main thread and I/O threads around <a href="https://github.com/valkey-io/valkey/pull/3324">lock-free queues</a>, credited in the release notes with an 8-17% throughput gain.</p>
<p>Based on these changes, we expected 8.0 to lift reads and writes for smaller items but still struggle to fill the network link for larger items. We expected 9.0 to bring GETs to line rate and 9.1 to improve writes.</p>
<p>This is what open-source competition buys everyone, including teams that never leave Redis. A performance design that sat for seven months as an unassigned issue became table stakes for both projects within two years of being filed.</p>
<h2 id="show-me-the-numbers">Show me the numbers</h2>
<p>We swept values from 1 MB to 64 MB across four Valkey releases on two nodes with 200 Gbps of bandwidth between them. The full setup is below.</p>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/8mb-release-journey.svg" alt="GET and SET bandwidth for 8 MB values across Valkey 7.2, 8.1, 9.0, and 9.1"></p>
<h3 id="valkey-72-reads-capped-at-35-gbps-writes-at-55">Valkey 7.2: reads capped at 35 Gbps, writes at 55</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-7-2-bandwidth.svg" alt="GET and SET bandwidth by value size for Valkey 7.2.13, with GET shown as a solid line and SET shown as a dashed line"></p>
<p>An 8 MB value delivers 21 Gbps on reads and 23 Gbps on writes, with p99 latencies of 174 and 164 ms. Turning on <code>io-threads-do-reads</code> moved only the 1 MB read cell.</p>
<h3 id="valkey-80-and-81-writes-take-off-large-reads-stay-put">Valkey 8.0 and 8.1: writes take off, large reads stay put</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-8-1-bandwidth.svg" alt="GET and SET bandwidth by value size, comparing Valkey 7.2.13 in red with Valkey 8.1.8 in orange"></p>
<p>The threading rebuild lifts our 8 MB write from 23 to 137 Gbps, six times faster, and 1 MB reads reach 183 Gbps. Larger reads settle at 30-33 Gbps whether the value is 8 MB or 64 MB. That flat floor points to a serial, per-byte bottleneck. Valkey 8.0.9 and 8.1.8 measured the same at every size, so one line carries both.</p>
<h3 id="valkey-90-large-gets-jump-to-line-rate">Valkey 9.0: large GETs jump to line rate</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-9-0-get-bandwidth.svg" alt="GET bandwidth by value size, comparing Valkey 8.1.8 in orange with Valkey 9.0.4 in green"></p>
<p>Every size from 1 MB to 64 MB reads at 190-201 Gbps. The 8 MB p99 falls from 112 to 24 ms, and a 64 MB read drops from just under a second to 230 ms. Writes do not move because the ingest path was never copy-bound. That ceiling waits for 9.1.</p>
<h3 id="valkey-91-more-headroom-for-writes">Valkey 9.1: more headroom for writes</h3>
<p><img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/release-9-1-set-bandwidth.svg" alt="SET bandwidth by value size, comparing Valkey 9.0.4 in dark green with Valkey 9.1.0 in light green"></p>
<p>With reads pinned at the network limit, the gain surfaces on writes. The 8 MB write rises from 134 to 166 Gbps, a 24% gain, while values from 12 MB to 64 MB gain 11-19%.</p>
<h2 id="the-short-run-held">The short run held</h2>
<p>To make sure the 15-second runs were not catching a lucky window, I reran every 8 MB GET and SET cell for 15 minutes. Valkey 9.1 is a good example: GET held 200.8 Gbps and SET held 165.5 Gbps, right on top of the original 201 and 166 Gbps results. Nothing sagged as the runs went on.</p>
<picture>
  <source media="(max-width: 640px)" srcset="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/long-run-stability-mobile.svg">
  <img src="https://www.gomomento.com/assets/content/blog/consistency-compounds-valkeys-journey-to-200-gbps/long-run-stability.svg" alt="Valkey 9.1 GET and SET throughput holding steady over 15 minutes" loading="lazy">
</picture>
<h2 id="detailed-results">Detailed results</h2>
<h3 id="reads-get">Reads (GET)</h3>
<p>GET-only, 32 connections, 100% hit rate.</p>
<div><table><caption>GET bandwidth (Gbps)</caption><thead><tr><th scope="col">Value size</th><th align="right" scope="col">7.2.13</th><th align="right" scope="col">8.1.8</th><th align="right" scope="col">9.0.4</th></tr></thead><tbody><tr><th scope="row">1 MB</th><td align="right">32</td><td align="right">183</td><td align="right">200</td></tr><tr><th scope="row">2 MB</th><td align="right">35</td><td align="right">55</td><td align="right">193</td></tr><tr><th scope="row">4 MB</th><td align="right">26</td><td align="right">45</td><td align="right">197</td></tr><tr><th scope="row">8 MB</th><td align="right">21</td><td align="right">33</td><td align="right">200</td></tr><tr><th scope="row">12 MB</th><td align="right">21</td><td align="right">32</td><td align="right">191</td></tr><tr><th scope="row">16 MB</th><td align="right">21</td><td align="right">32</td><td align="right">201</td></tr><tr><th scope="row">32 MB</th><td align="right">21</td><td align="right">32</td><td align="right">196</td></tr><tr><th scope="row">64 MB</th><td align="right">22</td><td align="right">31</td><td align="right">191</td></tr></tbody></table><div>
<p>Valkey 9.1 reads also hold line rate, so the table stops at 9.0.</p>
</div></div>
<div><table><caption>GET p99 latency (ms)</caption><thead><tr><th scope="col">Value size</th><th align="right" scope="col">7.2.13</th><th align="right" scope="col">8.1.8</th><th align="right" scope="col">9.0.4</th></tr></thead><tbody><tr><th scope="row">1 MB</th><td align="right">13</td><td align="right">2.2</td><td align="right">2.4</td></tr><tr><th scope="row">2 MB</th><td align="right">20</td><td align="right">19</td><td align="right">5.1</td></tr><tr><th scope="row">4 MB</th><td align="right">58</td><td align="right">41</td><td align="right">9.4</td></tr><tr><th scope="row">8 MB</th><td align="right">174</td><td align="right">112</td><td align="right">24</td></tr><tr><th scope="row">12 MB</th><td align="right">244</td><td align="right">166</td><td align="right">56</td></tr><tr><th scope="row">16 MB</th><td align="right">329</td><td align="right">238</td><td align="right">56</td></tr><tr><th scope="row">32 MB</th><td align="right">531</td><td align="right">489</td><td align="right">116</td></tr><tr><th scope="row">64 MB</th><td align="right">948</td><td align="right">1,020</td><td align="right">230</td></tr></tbody></table></div>
<h3 id="writes-set">Writes (SET)</h3>
<p>SET-only, 32 connections.</p>
<div><table><caption>SET bandwidth (Gbps)</caption><thead><tr><th scope="col">Value size</th><th align="right" scope="col">7.2.13</th><th align="right" scope="col">8.1.8</th><th align="right" scope="col">9.0.4</th><th align="right" scope="col">9.1.0</th></tr></thead><tbody><tr><th scope="row">1 MB</th><td align="right">51</td><td align="right">201</td><td align="right">201</td><td align="right">201</td></tr><tr><th scope="row">2 MB</th><td align="right">53</td><td align="right">200</td><td align="right">200</td><td align="right">201</td></tr><tr><th scope="row">4 MB</th><td align="right">55</td><td align="right">200</td><td align="right">200</td><td align="right">201</td></tr><tr><th scope="row">8 MB</th><td align="right">23</td><td align="right">137</td><td align="right">134</td><td align="right">166</td></tr><tr><th scope="row">12 MB</th><td align="right">23</td><td align="right">135</td><td align="right">139</td><td align="right">166</td></tr><tr><th scope="row">16 MB</th><td align="right">24</td><td align="right">137</td><td align="right">139</td><td align="right">165</td></tr><tr><th scope="row">32 MB</th><td align="right">24</td><td align="right">139</td><td align="right">134</td><td align="right">159</td></tr><tr><th scope="row">64 MB</th><td align="right">24</td><td align="right">130</td><td align="right">134</td><td align="right">149</td></tr></tbody></table></div>
<div><table><caption>SET p99 latency (ms)</caption><thead><tr><th scope="col">Value size</th><th align="right" scope="col">7.2.13</th><th align="right" scope="col">8.1.8</th><th align="right" scope="col">9.0.4</th><th align="right" scope="col">9.1.0</th></tr></thead><tbody><tr><th scope="row">1 MB</th><td align="right">8.7</td><td align="right">3.3</td><td align="right">3.2</td><td align="right">3.2</td></tr><tr><th scope="row">2 MB</th><td align="right">17</td><td align="right">6.5</td><td align="right">6.7</td><td align="right">5.3</td></tr><tr><th scope="row">4 MB</th><td align="right">35</td><td align="right">12</td><td align="right">13</td><td align="right">14</td></tr><tr><th scope="row">8 MB</th><td align="right">164</td><td align="right">41</td><td align="right">49</td><td align="right">34</td></tr><tr><th scope="row">12 MB</th><td align="right">289</td><td align="right">62</td><td align="right">57</td><td align="right">52</td></tr><tr><th scope="row">16 MB</th><td align="right">357</td><td align="right">77</td><td align="right">70</td><td align="right">60</td></tr><tr><th scope="row">32 MB</th><td align="right">1,580</td><td align="right">118</td><td align="right">143</td><td align="right">107</td></tr><tr><th scope="row">64 MB</th><td align="right">2,580</td><td align="right">237</td><td align="right">213</td><td align="right">206</td></tr></tbody></table></div>
<h2 id="the-setup">The setup</h2>
<p><strong>Machines.</strong> Two <code>c8gn.16xlarge</code> instances with Graviton4 and 200 Gbps networking in the same availability zone and cluster placement group, running Amazon Linux 2023.</p>
<p><strong>Server.</strong> We tested the official <code>valkey/valkey</code> Docker image at versions <code>7.2.13</code>, <code>8.1.8</code>, <code>9.0.4</code>, and <code>9.1.0</code>. We also measured <code>8.0.9</code>. It matched <code>8.1.8</code> within run-to-run noise at every size, so the tables show 8.1 as the 8.x column. Valkey 7.2 ran with the same flags as the newer versions. A control with <code>io-threads-do-reads yes</code> changed only the 1 MB GET cell, from 32 to 41 Gbps. Each version ran in a fresh container with host networking:</p>
<pre><code class="language-sh">docker run --network host --cpuset-cpus 8-23 \
  --ulimit nofile=32768:65536 valkey/valkey:&#x3C;version> \
  --save '' --appendonly no --io-threads 16 \
  --protected-mode no --maxmemory 60gb
</code></pre>
<p>Persistence was off. The 16 threads in Valkey’s <code>io-threads</code> count were one main thread plus 15 I/O workers. The process was pinned to cores 8-23 so it never fought the kernel for the cores doing network interrupt work.</p>
<p><strong>Interrupts.</strong> <code>irqbalance</code> was off on both machines. The ENA NIC was configured with four combined queues and their IRQs pinned to cores 0-3. Without this step, results wander from run to run as the kernel shuffles interrupts onto whatever cores the server or client threads happen to be using. If you benchmark at these speeds, pin your IRQs first and thank yourself later.</p>
<pre><code class="language-sh"># Run on both machines. ens50 is the ENA interface name on these instances.
sudo systemctl stop irqbalance
sudo ethtool -L ens50 combined 4
i=0
for irq in $(grep ens50 /proc/interrupts | awk '{print $1}' | tr -d ':'); do
  echo $((i%4)) | sudo tee /proc/irq/$irq/smp_affinity_list > /dev/null
  i=$((i+1))
done
</code></pre>
<p><strong>Client.</strong> We used <a href="https://github.com/cachecannon/cachecannon">valkey-lab</a>, built on cachecannon, with 16 worker threads pinned to cores 4-19, 32 connections, and pipeline depth 1. Reads and writes were measured in separate passes. Read passes prefilled the keyspace and ran at a 100% hit rate. Each short-run cell was a 15-second measurement after a five-second warmup, over a 500-key keyspace with 16-byte keys.</p>
<p>We kept concurrency at 32 connections because AWS caps a single TCP flow at roughly 9.5 Gbps. We verified 9.53 Gbps with iperf3. Saturating a 200 Gbps NIC requires spreading the load across flows.</p>
<p>One invocation per cell, with <code>-s</code> swept across the value sizes:</p>
<pre><code class="language-sh"># GET pass: prefill the 500-key keyspace, then measure at a 100% hit rate.
valkey-lab -h &#x3C;server> -c 32 -P 1 -t 16 --cpu-list 4-19 \
  -n 500 -r 100:0 --prefill --warmup 5s -d 15s \
  -s &#x3C;value_bytes> --key-size 16

# SET pass.
valkey-lab -h &#x3C;server> -c 32 -P 1 -t 16 --cpu-list 4-19 \
  -n 500 -r 0:100 --warmup 5s -d 15s \
  -s &#x3C;value_bytes> --key-size 16
</code></pre>
<p><strong>Scope of the result.</strong> This is a single-node, no-TLS, persistence-disabled test with 32 closed-loop connections, pipeline depth 1, and a 100% hit rate. Most table cells are one 15-second measurement after warmup. The results describe this rig and workload. Broader production claims need repeated runs and additional configurations.</p>
<p><em>We benchmarked whole releases rather than individual changes, so we treat the pull requests named in this post as leading explanations rather than proof of causality.</em></p>
<h2 id="what-this-means-if-you-run-valkey">What this means if you run Valkey</h2>
<p>Valkey 9.x delivers roughly 6x the 64 MB read bandwidth of 8.x while cutting p99 from just under a second to about 230 ms on this single-node, no-TLS test. Our earlier mixed-workload test also found far less collateral latency for small requests when a large read arrived.</p>
<p>The practical takeaway is narrower than “Valkey is always faster.” Valkey 8.1 is essentially flat on this workload, Valkey 9.0 changes the large-GET path, and Valkey 9.1 improves large SETs on this rig. If large values matter to your workload, test 9.x with your object-size distribution, concurrency, TLS, and persistence settings rather than extrapolating from a small-object benchmark, or from this one.</p>
<p>Valkey has consistently improved performance, memory efficiency, and availability at scale. Each version brings about a new set of improvements driven by issues faced by real users in production. A vibrant community where nobody is incentivized to withhold features for the sake of revenue and everyone is incentivized to chase continuous improvement is what makes Valkey truly special. Effort is additive. Consistency is multiplicative.</p>
<p>Serving megabyte-sized objects at wire speed is what I’ve been spending a lot of time on lately. If you are wrangling larger objects, let’s talk. <a href="https://valkey.io/slack/">Join me on Valkey Slack</a>.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
