<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Podo Stack]]></title><description><![CDATA[Tools that survived production. Weekly curation]]></description><link>https://podostack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!K687!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F647baa21-c6c0-4d23-bcf0-ddf3a7a641ed_500x500.png</url><title>Podo Stack</title><link>https://podostack.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 06 Aug 2026 03:54:54 GMT</lastBuildDate><atom:link href="https://podostack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Ilia]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[podostack@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[podostack@substack.com]]></itunes:email><itunes:name><![CDATA[Ilia Gusev]]></itunes:name></itunes:owner><itunes:author><![CDATA[Ilia Gusev]]></itunes:author><googleplay:owner><![CDATA[podostack@substack.com]]></googleplay:owner><googleplay:email><![CDATA[podostack@substack.com]]></googleplay:email><googleplay:author><![CDATA[Ilia Gusev]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[NetworkPolicy runs after DNAT: the hairpin nobody tests]]></title><description><![CDATA[The spec declares the ordering undefined, which is why the same manifest can allow traffic in staging and drop it in production]]></description><link>https://podostack.com/p/networkpolicy-after-service-dnat-hairpin</link><guid isPermaLink="false">https://podostack.com/p/networkpolicy-after-service-dnat-hairpin</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Wed, 05 Aug 2026 14:02:26 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!s_k0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!s_k0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!s_k0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!s_k0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!s_k0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!s_k0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!s_k0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!s_k0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!s_k0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!s_k0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!s_k0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98bcdb35-292f-4b1e-8dd1-d45061e3411b_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The policy was the careful kind. Egress allowed to <code>0.0.0.0/0</code> on 443, with the RFC1918 ranges carved out, so pods could reach the internet but not wander around the private network. It passed review. It worked in staging for three weeks.</p><p>In production it dropped a call that had nothing to do with private networks: a pod talking to the company's own API over its public hostname.</p><p>The manifest was identical in both places. What differed was where the packet got rewritten.</p><h2>The bit the manifest doesn't tell you</h2><p>A NetworkPolicy is written against addresses. Packets get matched against addresses. In between sits the part nobody diagrams: kube-proxy, or your CNI's replacement for it, rewriting the destination.</p><p>Pod sends to <code>203.0.113.10:443</code>, the load balancer address for <code>api.example.com</code>. That address resolves back into the cluster, so the packet never leaves. A Service picks it up and DNATs the destination to a backend pod, <code>10.244.3.17:8443</code>.</p><p>Now the policy gets evaluated. So what address does it see?</p><p>If enforcement happens before DNAT, it sees <code>203.0.113.10</code> and the rule matches <code>0.0.0.0/0</code>. Allowed.</p><p>If enforcement happens after DNAT, it sees <code>10.244.3.17</code>, which lands squarely in the <code>10.0.0.0/8</code> block you carved out. Dropped.</p><p>Same manifest. Same intent. Opposite outcome, decided by a mechanism the manifest can't see.</p><h2>The spec says it's undefined</h2><p>This is where it stops being a CNI bug and starts being a design gap. From the Kubernetes NetworkPolicy documentation:</p><p>&gt; Cluster ingress and egress mechanisms often require rewriting the source or destination IP of packets. In cases where this happens, it is not defined whether this happens before or after NetworkPolicy processing, and the behavior may be different for different combinations of network plugin, cloud provider, Service implementation, etc.</p><p>And more directly:</p><p>&gt; Connections from pods to Service IPs that get rewritten to cluster-external IPs may or may not be subject to ipBlock-based policies.</p><p>"May or may not" is doing a lot of work in a security primitive. Nobody's going to fix this in a patch release, because there is nothing to fix. Two implementations can both be conformant and disagree.</p><p>GKE takes the honest route and refuses the question: in Dataplane V2 you can't put a Pod or Service IP in <code>ipBlock.cidr</code> at all. The API rejects it rather than accepting a rule whose meaning depends on packet-rewrite ordering.</p><h2>Why ipBlock is the fragile selector</h2><p>Every other NetworkPolicy selector matches on something Kubernetes owns. <code>podSelector</code> matches labels. <code>namespaceSelector</code> matches labels. Those are stable facts in etcd, and no datapath component rewrites them mid-flight.</p><p><code>ipBlock</code> matches an IP address, which is the one field in the packet that half your infrastructure exists to change. kube-proxy rewrites it. The CNI rewrites it. The cloud load balancer rewrites it. Your egress gateway rewrites it. The rule's written against the one value that's explicitly in motion.</p><p>Which gives a rule of thumb worth more than any specific workaround: <strong>use label selectors for anything inside the cluster, and treat `ipBlock` as a tool for addresses that live outside the cluster.</strong> If the target has a Pod or a Service in front of it, select the Pod.</p><p>The carve-out policy above should have been two rules. Allow egress to the API by <code>podSelector</code> plus <code>namespaceSelector</code>, since it is in the cluster. Allow egress to <code>0.0.0.0/0</code> on 443 with private ranges excluded for everything that really is outside. Once the in-cluster case is handled by labels, the DNAT ordering stops mattering, because no rewritten address is ever matched against a CIDR.</p><h2>Where the dev/prod split comes from</h2><p>The reason this passes staging is that hairpin routing is not universal.</p><p>In a small cluster, a pod reaching for the external hostname of a Service in the same cluster often gets hairpinned by the node: the packet turns around at the local datapath and gets DNATed there. In a production setup with a real cloud load balancer, the same request can leave the node, reach the LB, and come back with a rewrite performed somewhere else entirely.</p><p>Two paths, two rewrite points, one policy. The environment where you test the policy is the environment where its behaviour is defined, and only there.</p><p>Testing that catches it looks like this: from a pod in the restricted namespace, curl the target the way the application actually addresses it. Not the Service DNS name if the app uses the public hostname. Not the public hostname if the app uses the Service. The address the app uses is the only one whose rewrite path you're testing.</p><pre><code>kubectl -n restricted run probe --rm -it --image=curlimages/curl --restart=Never -- \
  curl -sS -o /dev/null -w '%{http_code} %{remote_ip}\n' --max-time 5 https://api.example.com</code></pre><p><code>%{remote_ip}</code> is the useful part. It tells you which address curl actually connected to, which tells you whether you're on the hairpin path or the external one before you start reasoning about policy.</p><h2>The limits worth knowing</h2><p>Knowing your CNI's enforcement point helps, but it isn't a general answer. Cilium enforces in eBPF at the socket and tc layers, Calico in iptables or eBPF depending on the mode, and both move that boundary between versions. Anything you learn today is a fact about the version you're on, not about the API.</p><p>There's one thing the spec does guarantee, and it surprises people the other way: traffic to and from the node a pod runs on is always allowed, whatever your rules say. A policy that looks like it blocks the node isn't blocking the node.</p><p>And the honest summary of the whole area: NetworkPolicy is a good tool for expressing which workloads may talk to which workloads. It's a poor tool for expressing which IP ranges a workload may reach, because IP ranges aren't what it's built on. When you find yourself reaching for <code>ipBlock</code> to describe something inside your own cluster, that's usually the signal to go back to labels.</p><h3>Links</h3><ul><li><p><a href="https://kubernetes.io/docs/concepts/services-networking/network-policies/">Network Policies (Kubernetes docs)</a></p></li><li><p><a href="https://cloud.google.com/kubernetes-engine/docs/how-to/network-policy">GKE Dataplane V2 network policy limitations</a></p></li><li><p><a href="https://github.com/kubernetes/kubernetes/issues/114369">NetworkPolicy tests for north/south traffic (kubernetes#114369)</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Issue #029 - Delete your CPU limits: the CFS quota behind p99]]></title><description><![CDATA[How a limit becomes a 100ms budget, why threads spend it in parallel, and the two conditions that make removing it safe]]></description><link>https://podostack.com/p/issue-029-cpu-limits-cfs-quota-throttling</link><guid isPermaLink="false">https://podostack.com/p/issue-029-cpu-limits-cfs-quota-throttling</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Tue, 04 Aug 2026 14:01:44 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!_C1G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_C1G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_C1G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!_C1G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!_C1G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!_C1G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_C1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_C1G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!_C1G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!_C1G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!_C1G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0bbba30c-b4fc-46c6-a13c-9408ceca4247_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The service was using 340 millicores. Its limit was 1000. The dashboard showed 34% CPU utilisation, comfortably green, and it was being throttled 31% of periods. Somebody had already opened a ticket asking why we were paying for CPU we weren't using, and somebody else had opened one asking why p99 had tripled, and for two weeks nobody connected them.</p><p>They are the same ticket. The number on the dashboard is an average over a minute or a scrape interval. The kernel isn't averaging over a minute. It's making a decision every hundred milliseconds, and at that resolution the workload looks nothing like 34%.</p><p>Back in <a href="https://podostack.com/p/cold-start-pod-first-60-seconds-cgroup-stargz">issue #015</a> I went through what <code>requests</code> does on the way to the kernel - the conversion into <code>cpu.weight</code>, and the priority that got silently rescaled in the cgroup v1 to v2 migration. This is the other half of the pair, and it behaves nothing like it. <code>requests</code> is a relative claim that only matters when the node is contended. <code>limits</code> is an absolute budget that's enforced whether or not anybody else wants the CPU.</p><h2>&#127959;&#65039; Architectural Pattern: a limit is a budget, not a speed cap</h2><h3>What the number becomes</h3><p><code>limits.cpu: 1</code> doesn't tell the scheduler to run your container at most one CPU fast. There's no such mechanism. What it does is set two files.</p><p>On cgroup v2, one file:</p><pre><code># cpu.max
100000 100000</code></pre><p>Quota, then period, in microseconds. Your container may consume 100,000 microseconds of CPU time per 100,000 microsecond window. On cgroup v1 the same values live in <code>cpu.cfs_quota_us</code> and <code>cpu.cfs_period_us</code>. The period defaults to 100ms and Kubernetes doesn't expose a way to change it per pod.</p><p>The enforcement is brutally simple. The kernel tracks how much CPU time the cgroup has consumed in the current period. When the budget hits zero, every thread in that cgroup is dequeued and stays off-CPU until the period rolls over. Not slowed. Stopped.</p><p>That distinction is the whole issue. If a limit were a speed cap, a container using 34% of it would never notice. Because it's a budget with a hard refill boundary, a container can spend its entire budget in the first fifteen milliseconds of a period and then sit dead for eighty-five, and the minute-average will still say 34%.</p><h3>Threads spend the budget in parallel</h3><p>Service applications get hit by this far harder than batch ones, and the reason is arithmetic.</p><p>Quota is CPU-time, summed across all threads. A single-threaded process with <code>limits.cpu: 1</code> consumes 100ms of quota over 100ms of wall clock, and never throttles, because the two rates match exactly. Give the same limit to a process with eight runnable threads on an eight-core node, and it consumes 100ms of quota in 12.5ms of wall clock. It then has 87.5ms of enforced silence before the next refill.</p><p>Nothing about that's misconfiguration. That's a JVM with a thread pool sized to the machine, a Go binary with <code>GOMAXPROCS</code> picked up from the node's core count, a Node process with its libuv pool, nginx with <code>worker_processes auto</code>. Runtimes size themselves to the hardware they can see, and what they see is the node, not the cgroup.</p><p>The measured utilisation stays low the whole time, because the container really is idle for most of every period. It just happens to be idle in the specific way that adds up to 87 milliseconds of tail latency on any request unlucky enough to arrive at the wrong moment.</p><p>The metric that shows this isn't CPU usage. It's:</p><pre><code>rate(container_cpu_cfs_throttled_periods_total[5m])
  / rate(container_cpu_cfs_periods_total[5m])</code></pre><p>Anything meaningfully above zero on a latency-sensitive service is a finding. I've never seen a team that watches this and also has a mysterious p99 problem.</p><h3>Links</h3><ul><li><p><a href="https://docs.kernel.org/scheduler/sched-bwc.html">CFS Bandwidth Control (kernel docs)</a></p></li><li><p><a href="https://kubernetes.io/docs/tasks/configure-pod-container/assign-cpu-resource/">Assign CPU Resources to Containers (Kubernetes docs)</a></p></li></ul><div><hr></div><h2>&#127386; The Showdown: the 2019 kernel fix vs the bug people still have</h2><p>There are two separate throttling problems and they get conflated constantly, which matters because one of them is fixed and the other one is physics.</p><h3>The one that got fixed</h3><p>In 2019 Dave Chiluk at Indeed chased down containers that were being throttled heavily while using a fraction of their quota - not the parallel-spend effect above, but genuine phantom throttling where the accounting itself was wrong.</p><p>The cause was slice expiration. The global quota pool handed out CPU time to per-CPU run queues in slices, and any slice a run queue didn't fully use expired at the end of the period instead of returning to the pool. A threaded application that briefly touched many cores would scatter small unusable remnants across all of them, and the global pool would run dry while most of the quota sat stranded. His fix removed the expiration of cpu-local slices, and his test case on an 80-CPU machine with a 10ms/100ms allocation improved by roughly 30x.</p><p>It landed in Linux 5.4, in late 2019. If you're on anything modern, you have it, and any blog post about CPU throttling that predates 5.4 is describing a different kernel than the one you're running.</p><h3>The one that is still there</h3><p>What 5.4 didn't change - what nothing can change, short of removing the quota - is the parallel-spend arithmetic. Eight threads still burn the budget eight times faster than wall clock. That isn't a bug, it's what a CPU-time budget means.</p><p>Linux 5.14 added a genuine mitigation: <code>cpu.cfs_burst_us</code>, which lets a cgroup bank unused quota from quiet periods and spend it during a spike, up to a configured ceiling. It directly targets the bursty-service case, and it works.</p><p>Kubernetes can't set it. The request to expose it is <a href="https://github.com/kubernetes/kubernetes/issues/104516">issue #104516</a>, it has been open since 2021, and as of v1.36 there's still no pod-spec field for it. The teams that use burst today do it out of band - Koordinator and Alibaba's ACK both reach past the pod spec and adjust cgroup parameters on the node directly, which works and also means the value in your manifest no longer describes what's enforced.</p><p>So the honest state of things in 2026: the kernel has a good answer, the orchestrator doesn't expose it, and your remaining lever is whether the quota exists at all.</p><h3>Links</h3><ul><li><p><a href="https://lwn.net/Articles/792268/">sched/fair: Fix low cpu usage with high throttling (LWN)</a></p></li><li><p><a href="https://github.com/kubernetes/kubernetes/issues/104516">Use Linux CFS burst to get rid of unnecessary CPU throttling (kubernetes#104516)</a></p></li><li><p><a href="https://koordinator.sh/docs/user-manuals/cpu-burst">Koordinator CPU Burst</a></p></li></ul><div><hr></div><h2>&#128110; The Policy: the two conditions for deleting a limit</h2><p>Remove <code>limits.cpu</code>, keep <code>requests.cpu</code>, and the container gets a weight instead of a budget. Under contention it gets its proportional share. When the node is quiet it uses whatever is spare. No <code>cpu.max</code>, no periods, no throttling, and the tail latency that came from the quota boundary goes away.</p><p>This isn't a fringe position. It's what the kernel's own scheduling model wants you to do, and it's the default advice from most people who have debugged this at depth. It's also not unconditionally safe, and the two conditions are specific.</p><p><strong>Condition one: requests are actually set, everywhere, and they're honest.</strong> Without a limit, <code>requests</code> is the only thing standing between your workload and a neighbour's. If half the pods on the node have no requests either, you haven't removed a budget, you've removed the last piece of resource management and you'll find out during the next traffic spike. Enforce it with a policy rather than a convention - a Kyverno rule requiring <code>requests.cpu</code> on every container is about six lines, and it's the actual prerequisite for everything else here.</p><p><strong>Condition two: you aren't relying on Guaranteed QoS.</strong> A pod is Guaranteed only when every container has limits equal to requests for both CPU and memory. Drop the CPU limit and the pod becomes Burstable. Most of the time nobody cares. It matters if you use the static CPU Manager policy - exclusive core pinning requires Guaranteed and integer CPUs, so removing the limit silently turns pinning off for that pod. It also shifts the pod's position in node-pressure eviction ranking. If neither applies to you, and for most web services neither does, the reclassification is cosmetic.</p><h3>The honest part</h3><p>Two cases where I keep the limit.</p><p>Untrusted or third-party workloads, where the point isn't efficiency but blast radius. A vendor sidecar that spins on a bug should hit a wall, and a weight won't give you one.</p><p>Chargeback environments where the limit is the billing contract. The throttling is the product working as sold, and arguing about tail latency with the platform team isn't going to change the invoice.</p><p>And one thing I wouldn't do: keep the limit but raise it "to be safe". It's a real improvement - a bigger budget throttles less - but it's also the option that leaves you the least informed, because the throttling never goes to zero and you never learn what the workload actually needs. Either the budget is doing a job you can name, or it shouldn't be there.</p><p>The rollout that works is boring. Pick the service with the worst throttle ratio, remove the limit on that one, watch p99 and node CPU for a week. The first one is an experiment. After the third, it's a policy.</p><h3>Links</h3><ul><li><p><a href="https://kubernetes.io/docs/tasks/configure-pod-qos/">Configure Quality of Service for Pods</a></p></li><li><p><a href="https://kubernetes.io/docs/tasks/administer-cluster/cpu-management-policies/">Control CPU Management Policies on the Node</a></p></li></ul><div><hr></div><h2>What the dashboard was never going to tell you</h2><p>The reason this survives in so many clusters is that every instrument points the wrong way. Utilisation looks fine. Requests and limits look responsibly configured, because somebody set them deliberately. The limit was probably added by a thoughtful engineer during a capacity review, and it did what they asked. The cost landed somewhere they weren't looking.</p><p>The one number that would have caught it in an afternoon is the throttle ratio, and almost nobody graphs it, because it isn't in the default dashboard of anything.</p><p>We put it on the platform overview next to CPU and memory. Two services lit up immediately, both of them ones nobody had complained about, both of them well under half their limit. The ticket that got filed was for the loudest case, and it wasn't the worst one.</p><p>Questions? Feedback? Reply to this email. I actually read them.</p><ul><li><p>Ilia</p></li></ul>]]></content:encoded></item><item><title><![CDATA[CronJob: the 100-missed-schedules bug was fixed in 1.23]]></title><description><![CDATA[Two things everyone still tells you about Kubernetes CronJobs stopped being true four years ago, and the advice they came with is now the risk]]></description><link>https://podostack.com/p/cronjob-100-missed-schedules-fixed-in-123</link><guid isPermaLink="false">https://podostack.com/p/cronjob-100-missed-schedules-fixed-in-123</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Fri, 31 Jul 2026 14:02:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!36vX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!36vX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!36vX!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!36vX!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!36vX!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!36vX!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!36vX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/72148315-6ebe-4730-a666-30a9f7827239_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!36vX!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!36vX!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!36vX!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!36vX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F72148315-6ebe-4730-a666-30a9f7827239_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Search for why a Kubernetes CronJob stopped running and you'll land on the same answer within two results. It missed more than 100 scheduled runs, the controller gave up permanently, and the fix is to set <code>startingDeadlineSeconds</code>. The posts are confident and recent.</p><p>The behaviour they describe was removed in Kubernetes 1.23, in December 2021.</p><p>I went and read the controller across five release branches to be sure, because I had repeated this advice myself. Here is what's actually in there.</p><h2>What the old controller did</h2><p>Through 1.22, <code>getRecentUnmetScheduleTimes</code> counted missed start times and bailed out hard once the count passed 100:</p><pre><code>return []time.Time{}, fmt.Errorf(
  "too many missed start time (&gt; 100). Set or decrease .spec.startingDeadlineSeconds or check clock skew")</code></pre><p>That's an error return, not a warning. The caller stopped, no Job was created, and <code>status.lastScheduleTime</code> never advanced. That last one is what made it permanent. Next sync, the controller measured the gap from the same stale timestamp, counted the same growing pile of misses, and errored again. The CronJob looked healthy in <code>kubectl get cronjobs</code>. It would never run again without human intervention.</p><p>That's a nasty bug and it deserved every blog post it got. The recommended mitigation was correct too: setting <code>startingDeadlineSeconds</code> narrowed the window the controller counted misses over, so the count stayed under 100 and the wedge never triggered.</p><h2>What replaced it</h2><p>The rewritten controller became the default in 1.21 and the only one from 1.23. In current code the same condition produces an event and nothing else:</p><pre><code>if missedSchedules == manyMissed {
    recorder.Eventf(cj, corev1.EventTypeWarning, "TooManyMissedTimes",
      "too many missed start times. Set or decrease .spec.startingDeadlineSeconds or check clock skew")
    logger.Info("too many missed times", "cronjob", klog.KObj(cj))
}
return mostRecentTime, err</code></pre><p>The message survived, which is why the myth has such a long half-life - people still see <code>TooManyMissedTimes</code> in their events and reasonably conclude the documented consequence followed. It didn't. <code>mostRecentTime</code> is returned, the Job gets created, scheduling continues.</p><p>The second stale fact travels with the first. The old controller ran <code>wait.Until(jm.syncAll, 10*time.Second, stopCh)</code> - a flat ten-second poll across every CronJob in the cluster. That's where "don't set <code>startingDeadlineSeconds</code> below 10 seconds or the controller will miss the window" comes from, and it was true. The new one uses a delaying work queue and requeues each CronJob at its own next scheduled time, with a 100 millisecond pad for NTP skew:</p><pre><code>nextScheduleDelta = 100 * time.Millisecond</code></pre><p>No polling interval to lose runs in.</p><h2>What actually bites now</h2><p>Three things, and <code>startingDeadlineSeconds</code> is one of them - not as a fix, as the cause.</p><p><strong>The deadline still skips runs, and now that's its only job.</strong> The check is live in the current controller:</p><pre><code>tooLate := false
if cronJob.Spec.StartingDeadlineSeconds != nil {
    tooLate = scheduledTime.Add(time.Second * time.Duration(*cronJob.Spec.StartingDeadlineSeconds)).Before(now)
}</code></pre><p>Past the deadline, you get a <code>MissSchedule</code> warning event and the run is dropped. With the wedge gone, a field that used to buy you protection now only buys you skipped executions during any control-plane hiccup longer than the value you picked. Copying <code>startingDeadlineSeconds: 30</code> from a 2020 answer onto a modern cluster is a straight downgrade.</p><p>There's a wrinkle worth knowing if you alert on events. The controller doesn't record the miss in status - there's a standing TODO in the source about it - so <code>lastScheduleTime</code> stays put and the same miss is re-detected and re-reported on later syncs. One skipped run can produce a stream of <code>MissSchedule</code> events.</p><p><strong>`concurrencyPolicy: Forbid` drops runs silently.</strong> This is the failure I see most often in practice and it gets a fraction of the attention. A job that normally takes two minutes hits a slow dependency and runs for twenty. Every scheduled start in that window is skipped, not queued. The CronJob is working exactly as configured. Nothing is failing. You're simply not getting the runs you think you're getting, and the only trace is in events, which are gone in an hour by default.</p><p><strong>`backoffLimit` fails the Job, not the CronJob.</strong> The claim that exhausting <code>backoffLimit</code> suspends future execution isn't something I can find in the controller, and the behaviour doesn't match it: the Job goes to <code>Failed</code> and future scheduled runs continue normally. The actual problem is quieter. Failed Jobs sit in the namespace until <code>failedJobsHistoryLimit</code> (default 1) garbage-collects them, so by the time somebody looks, the evidence of the earlier failures is gone. <code>suspend: true</code> on the CronJob spec is the field that stops scheduling, and only if someone set it.</p><h2>What to do instead</h2><p>Leave <code>startingDeadlineSeconds</code> unset unless you can name the thing you're protecting against. The honest cases are narrow: a job that's worse than useless if it starts late, like a report bound to a market close, or one whose late start would collide with a maintenance window.</p><p>Alert on staleness, not on failure. Failures are visible. The dangerous mode is a CronJob that stops producing runs without erroring, and the metric that catches all three paths above is one expression:</p><pre><code>time() - kube_cronjob_status_last_schedule_time &gt; &lt;2x your interval&gt;</code></pre><p>That covers the deadline skip, the <code>Forbid</code> skip, and someone leaving <code>suspend: true</code> after an incident. It doesn't care which one happened.</p><p>Raise <code>failedJobsHistoryLimit</code> to something like 5 before you need it. The default of 1 means a flapping job erases its own history, and you get exactly one artifact to debug from.</p><h2>The limits worth knowing</h2><p>If you're on a cluster older than 1.23 - which in the wild usually means a vendor appliance or something in a regulated environment on a long support contract - the original bug is real and the original advice applies. Check the version before deciding which set of rules you're living under.</p><p>And there's one condition that still errors out rather than warning: if two consecutive schedule times work out to less than a second apart, <code>mostRecentScheduleTime</code> returns <code>time difference between two schedules is less than 1 second</code>. That takes a malformed schedule to hit, but it's the one place where a CronJob really does stop dead.</p><p>The general lesson is cheaper than any of this. Controller behaviour is a moving target and blog posts aren't. When something in Kubernetes behaves differently from what every result on the first page says, the source is one <code>gh api</code> call away and it's dated.</p><h3>Links</h3><ul><li><p><a href="https://github.com/kubernetes/kubernetes/blob/master/pkg/controller/cronjob/utils.go">pkg/controller/cronjob/utils.go</a></p></li><li><p><a href="https://github.com/kubernetes/kubernetes/blob/master/pkg/controller/cronjob/cronjob_controllerv2.go">pkg/controller/cronjob/cronjob_controllerv2.go</a></p></li><li><p><a href="https://kubernetes.io/docs/concepts/workloads/controllers/cron-jobs/">CronJob concepts (Kubernetes docs)</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[nginx proxy_pass: the variable that costs you the upstream block]]></title><description><![CDATA[The standard fix for stale pod IPs trades one visible failure for three that never show up in your logs, and resolve makes it unnecessary]]></description><link>https://podostack.com/p/nginx-proxy-pass-variable-drops-keepalive</link><guid isPermaLink="false">https://podostack.com/p/nginx-proxy-pass-variable-drops-keepalive</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Wed, 29 Jul 2026 14:00:31 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Mqs2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Mqs2!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Mqs2!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!Mqs2!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!Mqs2!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!Mqs2!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Mqs2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Mqs2!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!Mqs2!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!Mqs2!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!Mqs2!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F613645d9-bd85-44d6-9bc0-b4bccf84c7bc_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The 502s came in bursts, always right after a deploy, always for about ninety seconds. nginx was proxying to a headless service, and it had resolved those pod IPs once at startup. The pods behind them had been replaced. nginx didn't know and had no reason to find out.</p><p>The fix everybody reaches for is one line, and it works:</p><pre><code>set $backend "app.default.svc.cluster.local";
proxy_pass http://$backend;</code></pre><p>The 502s stop. The bill arrives later.</p><h2>Why the one-liner works</h2><p>When <code>proxy_pass</code> takes a literal address, nginx resolves it once, at configuration load, and caches the result for the process lifetime. No TTL, no re-resolution, no amount of waiting will change it. Only a reload will.</p><p>Put a variable in there and the rule changes. From the <code>ngx_http_proxy_module</code> docs:</p><p>&gt; Parameter value can contain variables. In this case, if an address is specified as a domain name, the name is searched among the described server groups, and, if not found, is determined using a resolver.</p><p>Two branches, and which one you land on decides everything that follows.</p><p>If the resolved string matches the name of an <code>upstream</code> block you've declared, nginx uses that group - and you're back to static peers, because the group's addresses were resolved at load time. The workaround does nothing. People hit this, conclude that variables in <code>proxy_pass</code> don't work, and go looking for a plugin.</p><p>If it doesn't match any declared group, nginx hands the name to <code>resolver</code> and honours the DNS TTL. That's the branch that fixes the 502s, and it's the one you get by writing out a full service FQDN, since nobody names an upstream block <code>app.default.svc.cluster.local</code>.</p><h2>What the second branch costs</h2><p>You're no longer proxying to an upstream group. You're proxying to whatever a DNS answer produced on this request. Everything the upstream block was doing goes away with it, and none of it announces itself.</p><p>Keepalive is the expensive one. The <code>keepalive</code> directive lives inside <code>upstream</code>; with no group there's no connection cache, and you pay a fresh TCP handshake - plus a TLS handshake, if the hop is encrypted - per request. This gap got wider in March 2026: nginx 1.29.7 switched upstream connections to HTTP/1.1 with keepalive enabled by default, 32 connections per worker. Upstream blocks now get connection reuse with no configuration at all. The variable path still gets none, so an upgrade that improved everyone else's latency silently widened the distance between you and them.</p><p>Load balancing goes too. <code>least_conn</code>, <code>hash</code>, weights - all <code>upstream</code> directives. What you get instead is nginx trying resolved addresses in the order the resolver returned them, which is round-robin at best and sticky at worst, depending on how your DNS server orders records.</p><p>Passive health checking goes. <code>max_fails</code> and <code>fail_timeout</code> mark a peer down after repeated failures and stop sending it traffic. Without a group there's no peer state to mark, so a pod that's accepting connections and failing every request keeps getting its share until DNS stops returning it.</p><p>And URI handling changes shape. With a literal <code>proxy_pass</code> that has a path component, nginx rewrites the request URI relative to the <code>location</code> prefix. With a variable it doesn't, so a config that was silently relying on that rewrite starts sending a different path upstream than it used to.</p><h2>The directive that makes the trade unnecessary</h2><p><code>resolve</code> on an upstream <code>server</code> has existed since 1.5.12, but until November 2024 it was NGINX Plus only, which is why so much of the advice online predates it and never mentions it. It landed in open source in 1.27.3 (mainline), then 1.28.0 (stable, April 2025).</p><pre><code>upstream app {
    zone app 64k;
    server app.default.svc.cluster.local resolve;
    keepalive 32;
}

resolver 10.96.0.10 valid=10s ipv6=off;

location / {
    proxy_pass http://app;
}</code></pre><p>nginx re-resolves the name in the background on DNS TTL and rewrites the peer list in place, without a reload. You keep the group, so you keep keepalive, the balancing method, and <code>max_fails</code>. It also lets nginx start when the name doesn't resolve yet, marking the peer down instead of refusing to load - which removes the other classic failure, where a restart during a DNS blip takes the proxy down entirely.</p><p>Two requirements, and both are hard errors if you skip them. The group must be in shared memory, so <code>zone</code> is mandatory - peers live in one place where every worker sees updates. And <code>resolver</code> must be configured, in the <code>http</code> block or in the <code>upstream</code> block itself. In Kubernetes point it at the cluster DNS service address and set <code>valid=</code> explicitly rather than trusting the record TTL, which is often 30 seconds and will feel slow during a rollout.</p><h2>The general shape of this bug</h2><p>The pattern isn't an nginx thing. A reverse proxy has to decide when a name becomes an address, and every proxy offers a fast path that resolves once and a slow path that resolves continuously, with different feature sets hanging off each.</p><p>Envoy calls them <code>STRICT_DNS</code> and <code>LOGICAL_DNS</code>, and the difference isn't what most people assume from the names: <code>STRICT_DNS</code> keeps a connection pool per resolved address, <code>LOGICAL_DNS</code> keeps one logical host and only uses the first address returned, which makes it cheaper and much worse at spreading load. HAProxy needs <code>resolvers</code> plus a <code>server-template</code> for the same effect, and a plain <code>server</code> line with a hostname is resolved at startup exactly like nginx.</p><p>If you take one thing: when a proxy suddenly starts tracking DNS changes, check what it stopped doing in exchange. There's almost always something, and it's almost never in the same paragraph of the docs.</p><h2>The limits worth knowing</h2><p><code>resolve</code> follows DNS, and DNS isn't a health signal. A pod that's <code>Terminating</code> stays in the endpoints list until the endpoint controller removes it, and your resolver caches that answer for <code>valid=</code>. Connection draining still needs <code>preStop</code> and a sane <code>terminationGracePeriodSeconds</code> on the pod side.</p><p>For most Kubernetes services this whole problem is avoidable. A normal <code>ClusterIP</code> gives you one stable virtual IP that never changes, and kube-proxy or the CNI handles the fan-out. The stale-IP problem is specific to headless services, where you deliberately asked for pod IPs. If you can't articulate why you need <code>clusterIP: None</code>, the cheapest fix is to stop using it.</p><h3>Links</h3><ul><li><p><a href="https://nginx.org/en/docs/http/ngx_http_upstream_module.html#server">ngx_http_upstream_module: server ... resolve</a></p></li><li><p><a href="https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_pass">ngx_http_proxy_module: proxy_pass</a></p></li><li><p><a href="https://blog.nginx.org/blog/keep-alive-to-upstreams-is-now-default-in-nginx-1-29-7">Keep-alive to upstreams is now default in NGINX 1.29.7</a></p></li></ul>]]></content:encoded></item><item><title><![CDATA[Issue #028 - ingress-nginx: four months archived, and nothing has broken yet]]></title><description><![CDATA[What actually retired, why InGate died too, the five real migration targets, and the annotations that don't port]]></description><link>https://podostack.com/p/issue-028-ingress-nginx-eol-migration</link><guid isPermaLink="false">https://podostack.com/p/issue-028-ingress-nginx-eol-migration</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Tue, 28 Jul 2026 14:01:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!3O6L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3O6L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3O6L!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!3O6L!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!3O6L!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!3O6L!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3O6L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!3O6L!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!3O6L!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!3O6L!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!3O6L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F057f2f5b-090c-41b1-a416-829369b9bbfc_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The repo went read-only on 24 March. I know the exact date because I went and looked it up this week, after an audit script flagged a cluster I had forgotten about - a small internal one, three services, an ingress-nginx install that had been sitting there since before I joined. It was serving traffic. It had been serving traffic every day for the four months since the project that maintains it stopped existing.</p><p>Nothing broke. Nothing is going to break tomorrow either. An archived controller doesn't stop reconciling, it doesn't phone home and refuse to start, it just keeps doing exactly what it did on 23 March and will keep doing it until the day something else forces the issue. Which means the deadline everybody planned around was never a real deadline. It was a date after which the failure mode changed from "we get a patch" to "we don't", and that change is invisible from the outside.</p><p>So this isn't a migration guide written before the fact. It's one written 124 days after, when the urgency has drained out of the room and the thing is still running in a lot of clusters, including probably one of yours.</p><h2>&#127959;&#65039; Architectural Pattern: what retired, and what everyone thinks retired</h2><h3>The announcement was narrower than the panic</h3><p>SIG Network and the Security Response Committee announced the retirement on 12 November 2025. Best-effort maintenance would continue until March 2026, and after that there would be no releases, no bug fixes, and no security patches. The repository - <code>kubernetes/ingress-nginx</code>, 19.5k stars - was archived on 24 March 2026 and is read-only now.</p><p>What retired is one specific implementation. Four things that people routinely believe also retired didn't:</p><p>The Ingress API itself is fine. It's feature-frozen, has been for years, and <code>networking.k8s.io/v1</code> isn't going anywhere. Your <code>Ingress</code> manifests are still valid Kubernetes objects. They just need something to act on them.</p><p>NGINX the web server is fine. It's F5's product, it ships releases on its own cadence, and 1.29.7 in March 2026 flipped upstream keepalive on by default. Unrelated project, unrelated lifecycle.</p><p><code>nginx/kubernetes-ingress</code> is fine, and this is the one that trips people up most. F5 maintains a separate ingress controller with a confusingly similar name, 5k stars, released v5.5.4 this month. It isn't the archived project. If you're on that one you aren't affected at all, and a surprising number of teams can't tell you offhand which of the two they installed.</p><p>And the reason for all of it wasn't technical. The project had one or two people doing the work, on their own time, after hours, for a controller that Datadog's telemetry put in roughly half of all cloud native environments. That ratio held for years. It was never sustainable, and eventually the maintainers said so out loud instead of continuing to absorb it.</p><h3>The official successor died first</h3><p>The plan, back in November, was InGate: a new controller built jointly by the ingress-nginx maintainers and the Gateway API community, meant to be the blessed migration path. It never got there. The repo is <code>kubernetes-sigs/ingate</code>, its description now literally starts with <code>[EOL]</code>, and it's archived alongside the thing it was supposed to replace.</p><p>I want to sit on this for a second, because it's the most useful fact in the whole story and it gets skipped in most of the migration posts. The recommended path didn't exist. If you started planning in December against "we will move to InGate when it's ready," you spent six months waiting for a project that shipped nothing in that window, and you're now four months past the deadline with no plan at all. That isn't a hypothetical - it's the single most common reason I've heard this month for why a cluster is still on the archived controller.</p><p>The lesson generalises past this incident. When an upstream announces a retirement and names a successor in the same breath, those are two separate claims with two separate confidence levels. The retirement is a decision, and decisions get executed. The successor is a forecast, and forecasts made by burnt-out maintainers about volunteer capacity aren't worth much. Plan against the retirement, treat the successor as a maybe.</p><h3>Links</h3><ul><li><p><a href="https://www.kubernetes.dev/blog/2025/11/12/ingress-nginx-retirement/">Ingress NGINX Retirement: What You Need to Know</a></p></li><li><p><a href="https://www.kubernetes.io/blog/2026/01/29/ingress-nginx-statement/">Statement from the Kubernetes Steering and Security Response Committees</a></p></li><li><p><a href="https://github.com/kubernetes/ingress-nginx">kubernetes/ingress-nginx (archived)</a></p></li></ul><div><hr></div><h2>&#127386; The Showdown: five migration targets and what each one actually costs</h2><p>There's no drop-in replacement. The committees said so in the January statement, and everyone who has done the move confirms it. What you're choosing between is five different shapes of work.</p><p><strong>Traefik</strong> is the closest thing to a soft landing, and it's the biggest project in the space by a wide margin - 64k stars, v3.7.9 last week. It ships an ingress-nginx provider that translates a good chunk of the annotation surface, so a simple Ingress moves with little more than a class change. Coverage isn't complete, and the gaps are exactly where you'd expect: anything that reached into raw nginx config. ModSecurity lives in Traefik Hub, which is the commercial tier.</p><p><strong>Envoy Gateway</strong> is the Gateway API answer with the least ceremony. 2.9k stars, v1.8.3, active. You're rewriting <code>Ingress</code> into <code>Gateway</code> plus <code>HTTPRoute</code>, which is real work, but the resource model is cleaner and the project isn't trying to also be a service mesh. If you've decided Gateway API is where you're going anyway, this is the shortest route with the fewest opinions attached.</p><p><strong>Contour</strong> is CNCF-incubating, 3.9k stars, v1.33.5, and has been doing Envoy-behind-a-controller longer than most. Same Gateway API rewrite cost. Smaller community than Envoy Gateway now, which wasn't true two years ago.</p><p><strong>kgateway</strong> is the one that surprised me when I went looking at current numbers - 5.6k stars, v2.4.0, and moving fast. Envoy-based, Gateway API native, with an AI-gateway angle bolted on that you can ignore if you don't need it. Younger, so the "someone has already hit my edge case and blogged about it" factor is lower.</p><p><strong>HAProxy Ingress</strong> at 1.2k stars is the small one, and I keep it on the list for exactly one reason: it has real embedded ModSecurity with Core Rule Set support. If your migration is blocked on a WAF ruleset you can't rewrite, this is the option that doesn't force you to stand up a separate WAF tier.</p><h3>The benchmark nobody wants to talk about</h3><p>John Howard - an Istio maintainer, who says up front that this makes him biased - published a benchmark harness comparing seven Gateway API implementations across attached routes, route propagation time, route scale, and traffic performance. The findings that stuck with me weren't the throughput numbers. They were the behavioural ones: Traefik consolidating all Gateways into a shared component, which is an isolation problem, and accepting traffic before config is applied. Cilium bottlenecking on a single CPU core for traffic performance.</p><p>The meta-point is worth more than any individual number. Nearly all of these sit on Envoy, and they still behave completely differently, because the proxy isn't the product. The controller is. Conformance tests cover the base case and stop well short of the things that will page you.</p><p>Treat the numbers as a prompt to run your own, not as a ranking. The harness is public.</p><h3>Links</h3><ul><li><p><a href="https://github.com/howardjohn/gateway-api-bench">gateway-api-bench (howardjohn)</a></p></li><li><p><a href="https://github.com/kubernetes-sigs/gateway-api/releases">Gateway API v1.6.1</a></p></li><li><p><a href="https://github.com/envoyproxy/gateway">Envoy Gateway</a></p></li></ul><div><hr></div><h2>&#128110; The Policy: audit the annotations before you pick the controller</h2><p>Every migration write-up I read starts with "pick your target." That's backwards, and it's why so many of these stall in week three. You can't evaluate targets until you know what you're actually using, and almost nobody does, because the annotation surface accumulated over five years across teams that have since reorganised.</p><p>Start here instead:</p><pre><code>kubectl get ingress -A -o json | jq -r '
  .items[] | .metadata.annotations // {} | keys[]
' | grep '^nginx.ingress.kubernetes.io/' | sort | uniq -c | sort -rn</code></pre><p>That gives you a frequency-ranked list of every ingress-nginx annotation in the cluster. Most of it will be <code>rewrite-target</code>, <code>ssl-redirect</code>, <code>proxy-body-size</code> - portable everywhere, boring, fine. What you're hunting for is the tail.</p><p>Four categories don't port cleanly, and finding them late is what turns a two-week migration into a two-quarter one.</p><p><strong>Snippets.</strong> <code>configuration-snippet</code> and <code>server-snippet</code> inject raw nginx config. There's no equivalent in Envoy, because it isn't nginx - the config languages have no common ground. Every snippet is a bespoke rewrite, and some of them encode behaviour that nobody remembers requesting. Later ingress-nginx releases already restricted these by default, after the IngressNightmare CVE work, so you may have fewer live than you fear.</p><p><strong>ModSecurity.</strong> <code>enable-modsecurity</code> and friends. Either HAProxy Ingress, or Traefik Hub, or a WAF tier in front. There's no path where the rules move unchanged into an Envoy-based controller.</p><p><strong>Auth subrequest.</strong> <code>auth-url</code> and <code>auth-signin</code> have equivalents nearly everywhere, but the semantics differ in the details - which headers get forwarded, what happens on a non-2xx, how the redirect is built. This is the category that passes staging and fails in production, because the difference only shows up on the error path.</p><p><strong>Session affinity edge cases.</strong> Plain cookie affinity is universal. Custom cookie names combined with path scoping and <code>affinity-mode: persistent</code> is where implementations diverge, and the failure is a user getting silently rebalanced mid-session rather than an error anyone sees.</p><h3>The honest part</h3><p>Two weeks minimum on staging before you touch production. I would say four if auth subrequest showed up in your audit. Run both controllers side by side with a split at the DNS or load balancer layer if your setup allows it, because the rollback story for "we swapped the ingress controller" is otherwise "we swap it back and hope the annotations still mean what they meant."</p><p>And a thing I've changed my mind about: I used to tell people to go straight to Gateway API on the grounds that the Ingress API is frozen and you'll have to move eventually. I now think that's bad advice for a team under time pressure. Moving to Traefik on Ingress semantics buys you a maintained controller in days rather than months, and the Gateway API migration is still there afterwards, on your schedule, decoupled from the security problem. Two smaller moves beat one big one when the first move is the urgent one. Gateway API is the right destination. It doesn't have to be the next step.</p><h3>Links</h3><ul><li><p><a href="https://kubernetes.io/docs/concepts/services-networking/ingress-controllers/">Ingress controller list (Kubernetes docs)</a></p></li><li><p><a href="https://github.com/traefik/traefik">Traefik ingress-nginx migration provider</a></p></li><li><p><a href="https://github.com/jcmoraisjr/haproxy-ingress">HAProxy Ingress</a></p></li></ul><div><hr></div><h2>Four months of nothing</h2><p>The thing that makes this hard to prioritise is that the risk curve doesn't look like the deadline curve. On 24 March nothing changed operationally. The controller in that forgotten cluster of mine is running the same binary, serving the same traffic, with the same behaviour it had in February. The only thing that changed is what happens next time somebody finds a bug in it, and nobody can tell you when that's.</p><p>That's also why I wouldn't let this sit much longer. Not because of a date that has already passed, but because the audit takes an afternoon and the answer determines whether your migration is a week or a quarter. Run the annotation query. If the tail is empty, you have a boring week ahead of you and you should just do it. If ModSecurity or auth subrequest shows up, you've found out now instead of in the middle of an incident, which is the whole point.</p><p>I ran it on the forgotten cluster. Twelve annotations, all boring, one afternoon of work. The one I haven't run it on yet is the big one, and I'm fairly sure I know why I keep not getting to it.</p><p>Questions? Feedback? Reply to this email. I actually read them.</p><ul><li><p>Ilia</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Every subdomain you own is already public]]></title><description><![CDATA[Certificate Transparency logs, crt.sh, and what to do on the mornings the web UI answers 502]]></description><link>https://podostack.com/p/crtsh-certificate-transparency</link><guid isPermaLink="false">https://podostack.com/p/crtsh-certificate-transparency</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Fri, 24 Jul 2026 14:02:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WmfG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!WmfG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!WmfG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!WmfG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!WmfG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!WmfG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!WmfG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!WmfG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!WmfG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!WmfG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!WmfG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F74516825-e3c5-4753-9666-a7eb9787e8b2_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The internal tool was on <code>grafana-staging.example.com</code>, behind the VPN, no public DNS record, not linked from anywhere. Someone asked whether it was really invisible. It took about fifteen seconds to prove it wasn't: the hostname was sitting in a public log, along with every other name we'd ever requested a certificate for.</p><p>Nobody leaked it. The certificate authority published it, because that's the deal.</p><h2>The deal</h2><p>Certificate Transparency means every publicly trusted TLS certificate gets written to append-only public logs at issue time. Browsers enforce it - a certificate that isn't in the logs doesn't get trusted, so there's no opting out while still having a working site.</p><p>The reason it exists is sound. Before CT, a CA could mis-issue a certificate for your domain and nobody would find out until it was used against someone. Now the mis-issuance shows up in a public log within hours, and the affected party can see it.</p><p>The side effect is the part most teams never think about: the log includes the hostnames. Every internal-sounding name you put in a SAN field is published, permanently, to anyone who cares to read.</p><h2>What that means in practice</h2><p>Type a domain into crt.sh and you get its certificate history. Use <code>%</code> as a wildcard - <code>%.example.com</code> - and you get every subdomain that has ever appeared in a certificate, including the ones that stopped existing years ago.</p><p>For an attacker doing reconnaissance this is the first stop, and it's free. <code>vpn.</code>, <code>jenkins-old.</code>, <code>admin-staging.</code>, the hostname of the acquisition you haven't announced yet - all of it, sorted by date. It costs nothing and touches none of your infrastructure, so there's nothing to detect and nothing to rate-limit.</p><p>The defensive uses run on the same data. Watch the logs for certificates issued on names that look like yours and you catch phishing infrastructure during setup rather than during the incident - <code>example-login.com</code> getting a certificate is someone building the page they intend to send your customers to. Watch your own domains and you catch the certificate a team issued outside the process, from a CA you don't use.</p><h2>When the front page is down</h2><p>Here's the operational bit, and it's the reason I'm writing this one now: crt.sh is a single popular free service with a very large database, and it falls over regularly. On the morning I wrote this, both the web UI and the JSON endpoint were returning 502.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!QDIh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!QDIh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 424w, https://substackcdn.com/image/fetch/$s_!QDIh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 848w, https://substackcdn.com/image/fetch/$s_!QDIh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 1272w, https://substackcdn.com/image/fetch/$s_!QDIh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!QDIh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png" width="1952" height="1546" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1546,&quot;width&quot;:1952,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!QDIh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 424w, https://substackcdn.com/image/fetch/$s_!QDIh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 848w, https://substackcdn.com/image/fetch/$s_!QDIh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 1272w, https://substackcdn.com/image/fetch/$s_!QDIh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5e9d35d0-45f2-41fd-b171-386b9eeaa5c1_1952x1546.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two things still work when that happens.</p><p>The database is separately reachable. crt.sh exposes read-only Postgres, and it was answering while the web tier returned 502:</p><pre><code>psql "postgresql://guest@crt.sh:5432/certwatch"</code></pre><p>Now the part that will save you twenty minutes, because every guide on the internet still gets it wrong. The canonical example everyone copies queries a <code>certificate_identity</code> table. That table is gone, and the error message tells you why:</p><pre><code>ERROR:  Sorry, the "certificate_identity" table has been superseded
by a Full Text Search index on the "certificate" table.</code></pre><p>The working shape uses that index instead:</p><pre><code>SELECT c.id, x509_commonName(c.certificate)
  FROM certificate c
 WHERE plainto_tsquery('example.com') @@ identities(c.certificate)
 LIMIT 100;</code></pre><p>Run that against <code>kubernetes.io</code> and among the first rows you get <code>spartakus.k8s.io</code> - a telemetry project that was shut down years ago, its hostname still sitting in the log where anyone can read it. That's the whole argument of this issue in one row.</p><p>Queries here are a language rather than a URL, so you can ask what the web form can't: group by issuer, filter by validity window, match a pattern across several registered domains in one pass. Set a <code>statement_timeout</code> and keep them narrow - broad ones get cancelled under load, which I also found out the direct way.</p><p>The other option is a different reader of the same logs. Cert Spotter's API takes a domain and returns issuances as JSON:</p><pre><code>curl "https://api.certspotter.com/v1/issuances?domain=example.com\
&amp;include_subdomains=true&amp;expand=dns_names"</code></pre><p>The logs themselves are the source of truth. crt.sh, Cert Spotter, Censys - they're all just indexes over the same public data, so when one is down you switch rather than wait.</p><h2>The limits worth knowing</h2><p>There's a lag between issuance and the record showing up in an index - usually minutes, sometimes hours. So for monitoring, alerting off a poll of an index is fine; treating its absence as proof that no certificate exists is not.</p><p>Private PKI is invisible here, by definition. If you run an internal CA for internal names, none of that reaches the public logs - which is also the practical mitigation for the problem this whole issue describes.</p><p>And wildcard certificates hide names rather than reveal them: a certificate for <code>*.example.com</code> publishes exactly one string. Teams that switch to wildcards for this reason are trading one exposure for a bigger blast radius on key compromise, which is a real trade, not an obvious win.</p><h2>What I'd actually do</h2><p>Run the wildcard query against your own domains this week and read the list properly. It is, at minimum, an accurate inventory of hostnames you've ever certified - and in my experience it always contains something nobody remembers creating.</p><p>Then set up monitoring on names close to yours, so brand-lookalike certificates page someone at issue time rather than after the phishing run starts. Cert Spotter and Censys both do this on a schedule; the free tiers cover a small domain list.</p><p>And when you name the next internal host, remember the name is going in a public log the moment it gets a certificate. <code>admin-legacy-payments.example.com</code> tells a stranger quite a lot.</p><h3>Links</h3><ul><li><p><a href="https://crt.sh/">crt.sh</a></p></li><li><p><a href="https://sslmate.com/certspotter/api/">Cert Spotter API</a></p></li><li><p><a href="https://certificate.transparency.dev/">Certificate Transparency project</a></p></li></ul><p>Questions? Feedback? Reply to this email. I actually read them.</p><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[Higress: the Ingress controller that became an AI gateway]]></title><description><![CDATA[Envoy and Istio underneath, Wasm plugins instead of Lua, and a pivot to MCP hosting that happened in about a year]]></description><link>https://podostack.com/p/higress-ingress-to-ai-gateway</link><guid isPermaLink="false">https://podostack.com/p/higress-ingress-to-ai-gateway</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Wed, 22 Jul 2026 14:03:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!neVw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!neVw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!neVw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!neVw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!neVw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!neVw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!neVw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/196744f4-c890-4103-a25a-0490343a02fc_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!neVw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!neVw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!neVw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!neVw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F196744f4-c890-4103-a25a-0490343a02fc_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I went looking at Higress in March because I wanted one thing gone: the stack where an L7 load balancer feeds an Ingress controller which feeds an API gateway, three components, three configs, three teams to page. Higress collapsed all three into one process on top of Envoy, and that was the whole pitch.</p><p>I went back to it this month and barely recognised the front page. It now calls itself a production-grade AI gateway for agent development and LLM API management. Same project, same Envoy core, completely different sales pitch.</p><p>That gap is worth a few minutes, because it tells you something about where gateways are going - and because the original reason to look at it still holds.</p><h2>What it actually is</h2><p>An API gateway built on Envoy, with Istio's traffic-management model underneath, that runs as your Kubernetes Ingress controller. The important part of that sentence is the part that isn't there: you don't have to install a service mesh. You get Istio's config machinery for routing without sidecars everywhere and without the operational weight that stops most mesh rollouts halfway.</p><p>It speaks HTTP/3, gRPC and Dubbo, and it reads service registries directly - Nacos, ZooKeeper, Consul, Eureka - which matters if your services aren't all Kubernetes-native and you've been gluing registry-to-Ingress adapters together by hand.</p><h2>The Wasm decision</h2><p>The extension model is where Higress made its actual bet. Kong and APISIX extend through Lua. Higress extends through WebAssembly, so a plugin can be written in Go, Rust, C++ or TypeScript.</p><p>The language list gets quoted a lot, but it's the second-order effect that changed my mind: a Wasm plugin runs in a sandbox. A plugin that panics takes down a request, not the gateway. Anyone who has had a Lua filter take out an ingress tier at 3am understands why that's worth a rewrite.</p><p>Config applies without a reload, which is the other thing Nginx-based gateways make you plan around.</p><h2>The AI turn</h2><p>Here's what changed since March, and it's not cosmetic.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2oAL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2oAL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 424w, https://substackcdn.com/image/fetch/$s_!2oAL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 848w, https://substackcdn.com/image/fetch/$s_!2oAL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 1272w, https://substackcdn.com/image/fetch/$s_!2oAL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2oAL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png" width="1912" height="1904" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1904,&quot;width&quot;:1912,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2oAL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 424w, https://substackcdn.com/image/fetch/$s_!2oAL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 848w, https://substackcdn.com/image/fetch/$s_!2oAL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 1272w, https://substackcdn.com/image/fetch/$s_!2oAL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F19524d99-7a47-4694-ac50-c98505a08498_1912x1904.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Higress now sits in front of model providers the way it used to sit in front of microservices: protocol conversion across model APIs so you write one integration instead of one per vendor, fallback to a second model when the first one errors, semantic caching, and token-level accounting so LLM spend stops being a single opaque line on the invoice.</p><p>Then the part I keep thinking about: it hosts MCP servers, and it converts existing HTTP APIs into MCP servers. Your internal APIs become tools an agent can call, and the gateway handles auth, traffic scheduling, parameter mapping and audit on the way through.</p><p>Think about what that means operationally. Every organisation that lets agents call internal services will need exactly that layer - identity, rate limits, an audit trail of which agent called what. Most teams are currently building it as bespoke glue in whatever framework their AI team picked. Higress is arguing it belongs in the gateway you already run, next to your ingress rules, and that argument is a lot stronger than it sounds at first pass.</p><h2>Where it stands</h2><p>The project moved out of the <code>alibaba</code> GitHub org into its own <code>higress-group</code>, which is the kind of unglamorous governance change worth noticing - single-vendor projects that survive usually make that move, and the ones that don't usually stall. It's Apache 2.0, around 8.9k stars, v2.2.3 landed at the end of June, and commits are steady.</p><p>The honest caveat is gravity. The registry integrations, the docs, the reference deployments - the centre of mass sits in the Alibaba Cloud ecosystem. That's not disqualifying, and nothing here is locked to it, but you'll spend more time on the unlit path than you would with ingress-nginx, and your search results will be in Chinese more often than not.</p><h2>Would I put it in front of prod</h2><p>For a plain Kubernetes ingress with nothing special going on: no. That's a boring problem with boring solutions, and boring is the correct answer for the component every request passes through.</p><p>For a team that already runs Envoy and wants extensions in a real language, or one that's about to build the auth-and-audit layer between agents and internal APIs from scratch: yes, at least as far as a proof of concept. The second case is the interesting one, because the alternative isn't another gateway - it's a pile of bespoke middleware nobody wants to own two years from now.</p><p>I'd still run it beside the existing ingress before I'd run it instead of one. Gateways are the wrong place to find out you were an early adopter.</p><h3>Links</h3><ul><li><p><a href="https://github.com/higress-group/higress">GitHub: higress-group/higress</a></p></li><li><p><a href="https://higress.cn/en/">Higress docs</a></p></li><li><p><a href="https://gateway-api.sigs.k8s.io/">Kubernetes Gateway API</a></p></li></ul><p>Questions? Feedback? Reply to this email. I actually read them.</p><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[Issue #027 - Parca, Popeye, Trivy vs Snyk, and the guide nobody reads]]></title><description><![CDATA[eBPF profiling with no code changes, a cluster that gets a letter grade, a scanner against a platform, and hardening advice written by a spy agency]]></description><link>https://podostack.com/p/issue-027-parca-popeye-trivy-snyk-hardening</link><guid isPermaLink="false">https://podostack.com/p/issue-027-parca-popeye-trivy-snyk-hardening</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Tue, 21 Jul 2026 15:25:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Q34_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q34_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q34_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!Q34_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!Q34_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!Q34_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q34_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png" width="1200" height="675" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:675,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Q34_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 424w, https://substackcdn.com/image/fetch/$s_!Q34_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 848w, https://substackcdn.com/image/fetch/$s_!Q34_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 1272w, https://substackcdn.com/image/fetch/$s_!Q34_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a2ecd07-6a99-46cb-841e-0ac5ab7f155e_1200x675.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two grab bags in a row. I said last week I was breaking the deep-dive format on purpose, and then this month handed me four more things that belong together.</p><p>Because this time there is a spine. Every one of these four answers the same question from a different angle: what is your cluster actually doing, and how would you find out? A profiler that needs nothing from your code. An auditor that gives the whole cluster a letter grade. Two scanners that disagree about what a vulnerability even is. And a hardening guide written by the people whose day job is breaking into things.</p><p>Side B of the mixtape.</p><h2>&#128640; Sandbox Watch: Parca</h2><h3>What it is</h3><p>Continuous profiling for Kubernetes, built on eBPF. It samples what every process on every node is doing with CPU and memory, all the time, and keeps the history.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jzyD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jzyD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 424w, https://substackcdn.com/image/fetch/$s_!jzyD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 848w, https://substackcdn.com/image/fetch/$s_!jzyD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 1272w, https://substackcdn.com/image/fetch/$s_!jzyD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jzyD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png" width="1992" height="1098" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1098,&quot;width&quot;:1992,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!jzyD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 424w, https://substackcdn.com/image/fetch/$s_!jzyD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 848w, https://substackcdn.com/image/fetch/$s_!jzyD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 1272w, https://substackcdn.com/image/fetch/$s_!jzyD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1663c1-d20c-45bf-89c9-dfc61714dcb1_1992x1098.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Not "profiling" as in the thing you enable for twenty minutes when something is already on fire. Profiling as in metrics: always on, always recording, there when you need to look backwards.</p><h3>How it works</h3><p>The collector runs on the node and reads stacks from the kernel. No SDK in your application, no rebuild, no restart, no sidecar. Because it works below the language, one agent covers the mixed reality of an actual cluster - Go, Java, Python, Rust, C++, and the kernel itself. It unwinds stacks even for binaries built without frame pointers, which is most of what you're running whether you know it or not.</p><p>Storing this is the hard part. A cluster's worth of stack samples is enormous, so Parca stores profiles in a columnar database built for exactly that shape of data, which is where the compression comes from.</p><h3>Why I like it</h3><p>Differential profiling. You pick two windows - before the release and after it - and Parca shows you which function got more expensive. Not "CPU went up 12%", which any dashboard tells you. Which function.</p><p>We lost most of a week last autumn to a memory graph that stepped up after every deploy. The dashboards said the API pods were heavier. The traces agreed the slow endpoint was slower. Neither got us past "somewhere in the API pods" until somebody profiled a container by hand on a Thursday and found a cache that never evicted. Twenty minutes of differential profiling would have pointed straight at it.</p><h3>When to use it</h3><p>The overhead claim is under 1%, against the 10-30% a traditional in-process profiler costs you - that's the whole reason "continuous" is possible at all.</p><p>Honest context: the continuous-profiling market consolidated when Pyroscope went to Grafana Labs, and if your shop is already all-in on Grafana, that's the path of least resistance. Parca's argument is that it stayed the eBPF-native, Kubernetes-native option instead of becoming a feature of somebody's platform. It's Apache 2.0, v0.28.0 landed in May, and commits are still going in this month - so this is a live project, not a preserved one. Which of those two paths matters more to you is a question about your stack, not about the profiler.</p><h3>Links</h3><ul><li><p><a href="https://github.com/parca-dev/parca">GitHub: parca-dev/parca</a></p></li><li><p><a href="https://www.parca.dev/docs/overview/">Parca docs</a></p></li></ul><h2>&#128142; The Hidden Gem: Popeye</h2><h3>What it is</h3><p>A read-only scanner that walks your live cluster and hands back a report card. Letter grade, A through F, plus every finding that dragged it down.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Eage!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Eage!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 424w, https://substackcdn.com/image/fetch/$s_!Eage!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 848w, https://substackcdn.com/image/fetch/$s_!Eage!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 1272w, https://substackcdn.com/image/fetch/$s_!Eage!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Eage!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png" width="1992" height="2084" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:2084,&quot;width&quot;:1992,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Eage!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 424w, https://substackcdn.com/image/fetch/$s_!Eage!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 848w, https://substackcdn.com/image/fetch/$s_!Eage!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 1272w, https://substackcdn.com/image/fetch/$s_!Eage!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F53e44f1a-f3fc-477c-ad05-9e72576d602a_1992x2084.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>It ships from the same author as k9s, which tells you what to expect: a single binary, fast, opinionated, made by somebody who clearly spends their day in a terminal looking at clusters.</p><h3>How it works</h3><p>Point it at a kubeconfig and run it. It reads - only reads - and reports what it finds: Secrets and ConfigMaps nobody references, containers with no resource requests or limits, pods with no liveness or readiness probe, Services whose selector matches nothing, port mismatches between Service and Pod, containers running as root, deprecated API versions you're still calling.</p><p>Each check is a sanitizer with its own code, so you can mute the ones that don't apply to you in a config file instead of learning to ignore them.</p><h3>Why I like it</h3><p>The grade is the trick, and it's a psychological one rather than a technical one. A list of 340 warnings gets closed. A cluster with a <strong>D</strong> gets discussed in standup. Same data, completely different organizational outcome - and I've watched that happen more than once.</p><p>The second use is upgrade prep. Before a version bump, Popeye tells you what still calls APIs that are about to disappear, which is a lot cheaper to learn on a Tuesday than during the upgrade window.</p><h3>When to use it</h3><p>The caveat you need before you adopt it: Popeye is quiet. The last release, v0.22.1, shipped in January 2025, and the last commits landed at the end of that year. Meanwhile k9s - same author - put out a release last month. Draw your own conclusion about where the attention went.</p><p>I'm still recommending it, and here's the reasoning. This is a read-only tool you run manually against a cluster, not a controller you put in the critical path of anything. The blast radius of "unmaintained" here is that new checks stop arriving and new API deprecations eventually go unnoticed - annoying, not dangerous. That's a very different risk than running an unmaintained admission webhook. Just know what you're picking up, and don't wire it into a pipeline that has to keep working in 2028.</p><p>Right tool, right stage, too. Popeye reads the live cluster, so it catches what actually got deployed, including the thing somebody applied by hand in March. Static YAML analysis in CI catches problems before they land but can't see drift. An admission controller enforces rules but doesn't tell you about the mess you already have. Three different jobs; Popeye only does the third.</p><p>Run it read-only. Give it a ReadOnly RBAC role and nothing more - it never needs write access, so don't hand it any.</p><h3>Links</h3><ul><li><p><a href="https://github.com/derailed/popeye">GitHub: derailed/popeye</a></p></li><li><p><a href="https://popeyecli.io/">Popeye docs</a></p></li></ul><h2>&#9876;&#65039; The Showdown: Trivy vs Snyk</h2><h3>The comparison everyone gets wrong</h3><p>These get benchmarked head to head constantly, and the benchmarks are close to meaningless, because one is a tool and the other is a product. Trivy is a scanner: a binary, in your pipeline, exit code 1, done. Snyk is a vulnerability management platform that happens to contain a scanner.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ImZe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ImZe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 424w, https://substackcdn.com/image/fetch/$s_!ImZe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 848w, https://substackcdn.com/image/fetch/$s_!ImZe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 1272w, https://substackcdn.com/image/fetch/$s_!ImZe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ImZe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png" width="1992" height="1322" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1322,&quot;width&quot;:1992,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ImZe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 424w, https://substackcdn.com/image/fetch/$s_!ImZe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 848w, https://substackcdn.com/image/fetch/$s_!ImZe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 1272w, https://substackcdn.com/image/fetch/$s_!ImZe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6b3459fd-5ab9-4bbe-81b7-2a4b6dfef72b_1992x1322.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I've sat in the meeting where somebody put a "finds 14% more CVEs" slide on the wall. Nobody in the room asked what happened to the findings after they were found, which turned out to be the only question that mattered six months later.</p><h3>Where they actually differ</h3><p><strong>The database.</strong> Trivy aggregates public sources - NVD, and the per-distro feeds from Red Hat, Debian, Alpine. It's excellent on OS packages and rarely lies to you there. Snyk maintains its own curated database, which typically knows about things before NVD does and carries more context per finding.</p><p><strong>Reachability.</strong> This is Snyk's real weapon, and Trivy has no full answer to it. Trivy tells you a vulnerable version of a library is present. Snyk can tell you your code never calls the vulnerable function. On a large monolith that difference is the difference between a security team triaging 400 criticals and triaging 40.</p><p><strong>What you get at the end.</strong> Trivy gives you a report - JSON, SARIF, a table - and expects you to build the process around it. Snyk opens the pull request that bumps the dependency. If your goal is developers fixing their own findings without a security engineer chasing them, that gap is the entire product.</p><p><strong>Price and gravity.</strong> Trivy is Apache 2.0, unlimited, runs entirely on your infrastructure, and never sends your code anywhere - and it is very much alive, with v0.72.0 out at the end of June. Snyk is per-developer at enterprise scale and that math gets steep - and it wants your source. For some organizations that's disqualifying before the feature comparison starts.</p><h3>My read</h3><p>If you're building a platform, keep Trivy in the pipeline. It's the boring correct default: free, fast, self-hosted, no vendor in the critical path of your build.</p><p>Buy the platform when the bottleneck stops being detection and becomes triage - too many findings, too few people, developers who won't act on a JSON report. That's a real problem and it's worth money to solve. Just be honest that you're buying triage and developer experience, not better detection.</p><p>And note the shape of it, because this newsletter keeps running into the same shape: the open tool does the work, the commercial platform sells you the workflow around the work. Same as the last two issues, different corner of the stack.</p><h3>Links</h3><ul><li><p><a href="https://github.com/aquasecurity/trivy">GitHub: aquasecurity/trivy</a></p></li><li><p><a href="https://trivy.dev/">Trivy docs</a></p></li><li><p><a href="https://snyk.io/">Snyk</a></p></li></ul><h2>&#128110; The Policy: the hardening guide written by the NSA</h2><h3>Why this one is worth your time</h3><p>The NSA and CISA publish a Kubernetes Hardening Guide. It is free, it is specific, and almost nobody in our industry has read it - which is odd, given that it's a defensive checklist from the agency with the most credible offensive team on the planet.</p><p>I opened it expecting compliance boilerplate. What's in there reads much more like the ways clusters actually get taken, written by people who take them for a living.</p><h3>What it actually asks for</h3><p>Nothing exotic, which is the point:</p><p>Containers run as a non-root user, with a read-only root filesystem, dropping every capability they don't need. No privileged pods, no host namespaces, no hostPath mounts into anything that matters.</p><p>Network policy that denies by default. Not "we have some network policies" - a default-deny posture where a compromised pod can't reach the rest of the cluster, because flat pod networking is what turns one bad container into a cluster-wide incident.</p><p>RBAC that's actually least privilege, with cluster-admin as a rare exception rather than the shape of every service account. Audit logging turned on and shipped somewhere off-cluster, because logs on a compromised node are evidence you no longer control. Images scanned, and the control plane and etcd locked down.</p><h3>The honest part</h3><p>The guide came out in August 2021 and was last revised in August 2022. CISA now files the announcement under archived content, with the standard banner about information that may no longer reflect current policy. Four years is a long time in this ecosystem.</p><p>But read what's actually in it and almost none of it has expired, because the failure modes haven't changed. What has changed is that Kubernetes grew native answers: Pod Security Standards now cover a good chunk of the container-level requirements, and the <code>restricted</code> profile is close to a working implementation of that whole section. So treat the guide as the argument for *why*, and Pod Security Standards plus a policy engine as the *how*.</p><p>The part I actually took from it was smaller than the document. We picked four of those requirements, wrote them down as the standard for our platform, and moved them into admission. The wiki page that had said roughly the same thing had been sitting there ignored for about a year and a half before that.</p><p>Then check it from the outside - which is where this issue closes its own loop. Popeye will tell you which containers run as root and which workloads have no probes. Trivy will tell you which images shouldn't be there. The guide tells you what the answer is supposed to be.</p><h3>Links</h3><ul><li><p><a href="https://www.cisa.gov/news-events/alerts/2022/03/15/updated-kubernetes-hardening-guide">CISA: Updated Kubernetes Hardening Guide</a></p></li><li><p><a href="https://kubernetes.io/docs/concepts/security/pod-security-standards/">Kubernetes Pod Security Standards</a></p></li></ul><p>That's side B. Four tools that answer the same question - what is this cluster really doing - from four different distances.</p><p>Popeye is the one that'll surprise you fastest. It took about five minutes against our staging cluster and came back with a grade I chose not to mention in standup.</p><p>Questions? Feedback? Reply to this email. I actually read them.</p><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[OpenFeature: swap your feature-flag vendor in one line]]></title><description><![CDATA[Evaluation API, providers, hooks, LaunchDarkly migration, CNCF incubating spec]]></description><link>https://podostack.com/p/openfeature-vendor-swap</link><guid isPermaLink="false">https://podostack.com/p/openfeature-vendor-swap</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Fri, 17 Jul 2026 14:02:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K687!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F647baa21-c6c0-4d23-bcf0-ddf3a7a641ed_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The renewal quote landed in March and the flag vendor wanted roughly triple. Not for new capabilities - for the same targeting rules and the same dashboard, priced against a seat count that had grown while nobody watched. The team I was advising did the estimate everyone does at that point: how much to move off? Someone grepped the codebase. Eleven hundred call sites, every one of them importing the vendor's SDK directly, some of them consuming vendor-specific context objects three layers deep in business logic. The migration priced out at two quarters of platform work. They paid the invoice.</p><p>That grep result is the exact thing OpenFeature exists to prevent. I gave it one section in a grab-bag issue a while back - <a href="https://podostack.com/p/dapr-kargo-wasmedge-koordinator-openfeature">issue #012</a> - and the pitch back then was "OpenTelemetry, but for feature flags." That's still the right one-liner. What I want to do here is walk the actual architecture, because whether the promise holds depends entirely on where the seam sits.</p><h2>Three pieces, one seam</h2><p>The spec splits flag evaluation into three parts, and the split is the whole product.</p><p>The <strong>Evaluation API</strong> is what your code touches: <code>getBooleanValue("new-checkout", false, ctx)</code> and its friends for strings, numbers and objects. It belongs to OpenFeature, not to any vendor. Those eleven hundred call sites, written against this API, contain zero vendor imports.</p><p>The <strong>provider</strong> is the adapter that speaks one vendor's protocol - LaunchDarkly, Flagsmith, PostHog, a JSON file on disk, your own homegrown service. It's registered once, at startup. This line is the only place a vendor name appears in the application, which is why switching backends is a one-line diff plus a config migration, instead of two quarters.</p><p><strong>Hooks</strong> wrap every evaluation with before/after/error stages. Access logging, OpenTelemetry spans per flag check, validation that a flag actually exists before you bet a request on it - all of it lands in one place instead of being sprinkled across call sites. Teams underuse these badly; a hook that emits an event on every evaluation is how you find the flags nobody has toggled in a year.</p><p>The context object rounds it out: user key, region, plan tier, whatever your targeting rules segment on, passed in a standard shape so the provider translates it rather than your code adapting to the provider.</p><h2>What it doesn't do</h2><p>Fair's fair: OpenFeature standardizes evaluation, not management. There's no standard API for creating flags, wiring targeting rules or archiving stale ones - that stays in each vendor's dashboard and each vendor's Terraform provider, and moving those rules between systems is still real migration work. The one-line swap covers your application code, which is the part that used to be unmovable. The rules move by hand.</p><p>The ecosystem question also answers itself these days: the project is CNCF incubating, SDKs cover the mainstream server and client stacks, and the official provider registry lists 30-plus backends including every vendor whose renewal quote might one day surprise you. When the big flag vendors all maintain providers for the abstraction layer that makes leaving them easy, the standard has won the argument.</p><h2>The migration that convinced me</h2><p>The pattern I've now seen work: wrap the current vendor first, migrate nothing. Every new flag check goes through OpenFeature backed by the existing provider; old call sites get converted opportunistically, whenever a file is already open for other reasons. Six months later the vendor import count is near zero and the company has options it didn't have - including the boring option of staying, but negotiating with a grep result that says leaving costs a sprint instead of two quarters.</p><p>Flags are infrastructure with a pricing model attached. Infrastructure with a pricing model attached is exactly the place to keep an exit visible.</p><h3>Links</h3><ul><li><p><a href="https://openfeature.dev/">OpenFeature</a></p></li><li><p><a href="https://github.com/open-feature">GitHub: open-feature</a></p></li><li><p>
<a href="https://podostack.com/p/dapr-kargo-wasmedge-koordinator-openfeature">Issue #012: tools from the future</a>
</p></li></ul><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[Gateway API request mirroring: a shadow copy of prod traffic]]></title><description><![CDATA[HTTPRoute RequestMirror filter, fire-and-forget copies, idempotency traps, fraction-based sampling]]></description><link>https://podostack.com/p/gateway-api-request-mirroring</link><guid isPermaLink="false">https://podostack.com/p/gateway-api-request-mirroring</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Wed, 15 Jul 2026 14:02:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K687!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F647baa21-c6c0-4d23-bcf0-ddf3a7a641ed_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The rewrite had passed every test we had. Six months of porting a payment-adjacent service from its legacy runtime, a test suite everyone trusted, load tests green, staging green. The one thing nobody could answer was the only question that mattered: what happens when it meets the traffic we can't imagine - the malformed headers, the retries from that one mobile client, the requests that only show up on the last Friday of the month. Staging never has those. Prod has nothing but.</p><p>Request mirroring is the answer to that specific fear, and in Gateway API it's a first-class filter rather than a vendor annotation. The gateway takes each incoming request, sends it to your real backend as usual, and sends a copy to a shadow backend. The response from the shadow gets dropped on the floor. Users never see it, latency doesn't grow because the copy is fire-and-forget, and your v2 gets a full day of real production traffic without owning a single response.</p><h2>The filter itself</h2><p>Mirroring lives in <code>HTTPRoute</code>, in the <code>filters</code> list of a rule:</p><pre><code>apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: store-route
spec:
  parentRefs:
  - name: my-gateway
  hostnames:
  - "store.example.com"
  rules:
  - backendRefs:
    - name: store-v1
      port: 8080
    filters:
    - type: RequestMirror
      requestMirror:
        backendRef:
          name: store-v2-shadow
          port: 8080</code></pre><p>Traffic flows to <code>store-v1</code> and answers come from <code>store-v1</code>. The copy goes to <code>store-v2-shadow</code>, whose responses the gateway must ignore - that's in the spec, not an implementation courtesy. Because it's Gateway API rather than an Ingress annotation, the same manifest works whether the underlying data plane is Envoy Gateway, Istio, Cilium or anything else that implements the filter. RequestMirror sits at "extended" conformance, so check your implementation's support matrix once - I've yet to meet a major one that skips it.</p><p>Full traffic is not always what you want; doubling the load on day one rarely is. The spec has a <code>fraction</code> field for that - <code>numerator</code> over <code>denominator</code> (default 100) - so you can start by mirroring 1% and turn the dial as confidence grows.</p><h2>What it's for</h2><p>The rewrite scenario is the obvious one: run v2 in the shadow for a week, compare its error rates and latency histograms against v1 on identical input, and the argument about readiness settles itself. The same trick warms caches before a cutover - the shadow instance has hot caches on switch day because it's been serving phantom traffic all week. And for the bug that only reproduces in prod, a mirror to a debug deployment with verbose logging gets you the evidence without touching the pod that's actually serving users.</p><h2>Where it bites</h2><p>The trap that matters is side effects. A mirrored request looks exactly like a real one to the service receiving it. If your shadow v2 writes to the production database, sends emails or charges cards, congratulations - you've built a duplication engine. The shadow needs its own isolated database, stubbed external calls, or a strictly read-only path. I'd treat this as a hard gate: no mirror until someone has written down where every write in the shadow path lands.</p><p>Two smaller ones. The gateway pays for the cloning - at high traffic, mirroring doubles the data plane's work, so watch CPU on the gateway pods, not just the backends. And the mirrored request keeps the original <code>Host</code> header; if the shadow routes by hostname, you'll need a <code>RequestHeaderModifier</code> in the chain before things line up.</p><p>The last thing isn't a trap so much as a prerequisite: the gateway throws the shadow's responses away, so metrics, traces and logs are the only window into whether v2 actually behaved. Mirroring without observability on the shadow is just heating the datacenter.</p><h3>Links</h3><ul><li><p><a href="https://gateway-api.sigs.k8s.io/guides/http-request-mirroring/">Gateway API: HTTP request mirroring guide</a></p></li><li><p>
<a href="https://gateway-api.sigs.k8s.io/references/spec/">Gateway API spec reference</a>
</p></li></ul><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[Issue #026 - Timoni, Dolt, OpenBao, and the account that should never log in]]></title><description><![CDATA[CUE-typed Kubernetes packages, Prolly Trees, the MPL fork of Vault, zero-usage emergency accounts]]></description><link>https://podostack.com/p/issue-026-timoni-dolt-openbao-break-glass</link><guid isPermaLink="false">https://podostack.com/p/issue-026-timoni-dolt-openbao-break-glass</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Tue, 14 Jul 2026 14:03:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K687!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F647baa21-c6c0-4d23-bcf0-ddf3a7a641ed_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The last thirteen issues each went deep on one thing. This week I'm breaking that on purpose. Four tools crossed my desk this month that don't need three thousand words each - but every one of them changed how I think about a piece of my stack, and none of them would leave my head.</p><p>So, old-school grab bag: a package manager that finally makes Kubernetes configs type-checked, a MySQL you can branch like a repo, a fork that outlived the drama that created it, and an account whose entire job is to never log in.</p><div><hr></div><h2>&#128640; Sandbox Watch: Timoni</h2><h3>What it is</h3><p>A package manager for Kubernetes from Stefan Prodan - the Flux guy - that throws out Helm's Go-templates-in-YAML and replaces them with CUE, an actual configuration language with types.</p><p>If you've ever debugged a chart where <code>{{ indent 4 }}</code> was off by two spaces, you already know why this exists.</p><h3>How it works</h3><p>A Timoni module is CUE code: schema and values get unified, and if a value doesn't match the schema, the build fails on your laptop - not halfway through a deploy. Constraints live in the type itself. <code>port: int &amp; &gt;1024 &amp; &lt;65535</code> isn't a comment or a JSON schema in a separate file somebody forgot to update; it's the definition of the field.</p><p>Modules ship as OCI artifacts, so they sit in the same registry as your images, signed and versioned. Bundles describe whole stacks - which modules, what order, which clusters. And <code>timoni build</code> prints exactly what will hit the API server, which is more than I can say for some dry-runs I've trusted.</p><h3>Why I like it</h3><p>Garbage collection is built in: remove a resource from the config and Timoni deletes it from the cluster instead of leaving it orphaned. The diff output has syntax highlighting. Small things, but they're the things you touch daily.</p><h3>When to use it</h3><p>Honest caveat first: v0.27.0 shipped on July 2nd and the project is alive, but the README still warns that APIs may break before 1.0, and the module ecosystem is a rounding error next to Helm's charts. I wouldn't migrate a company on it. For an internal platform where you write every module yourself? That's exactly where I'd start.</p><h3>Links</h3><ul><li><p><a href="https://github.com/stefanprodan/timoni">GitHub: stefanprodan/timoni</a></p></li><li><p><a href="https://timoni.sh/">Timoni docs</a></p></li></ul><div><hr></div><h2>&#128142; The Hidden Gem: Dolt</h2><h3>What it is</h3><p>A MySQL-compatible database where tables have Git semantics. <code>dolt commit</code>, <code>dolt branch</code>, <code>dolt merge</code> - on rows, not files. Any MySQL client connects to it and has no idea anything unusual is underneath.</p><h3>How it works</h3><p>The trick is a structure called Prolly Trees - a cross between the B-trees databases use and the Merkle trees Git uses. Diffing two versions of a huge table doesn't scan the data; it compares subtree hashes and descends only where they differ. History is stored with structural sharing, so a new version keeps only what changed.</p><p>Merge conflicts get resolved at the cell level. Two analysts edit the same row in different branches, and Dolt shows you both values and asks.</p><h3>Why I like it</h3><p>The reproducible-ML story sold me. Pin the commit hash of your training data next to the commit hash of your model code, and "what exactly did this model learn from" stops being archaeology. Same shape works for config databases: changes go through a pull request on data, with a reviewable diff of rows.</p><h3>When to use it</h3><p>Writes run 2-4x slower than stock MySQL - Merkle hashing isn't free - so this is not your hot OLTP path. Reference data, feature stores, config, the datasets your agents keep rewriting: that's the sweet spot, and v2.1.10 landed in late June, so the project is moving fast.</p><h3>Links</h3><ul><li><p><a href="https://github.com/dolthub/dolt">GitHub: dolthub/dolt</a></p></li><li><p><a href="https://www.dolthub.com/docs/">Dolt docs</a></p></li></ul><div><hr></div><h2>&#9876;&#65039; The Showdown: OpenBao vs Vault</h2><h3>The story so far</h3><p>In August 2023 HashiCorp moved Vault from MPL to BUSL - source available, but no competing products allowed. The community forked Vault 1.14 into OpenBao under the Linux Foundation, IBM loudly backing it. You know this part.</p><p>Here's the part that still makes me grin: in February 2025, IBM bought HashiCorp for $6.4 billion. The company that bankrolled the open fork now owns the thing it forked from. Nobody has blinked - as of this summer the BUSL stands, and OpenBao keeps shipping under MPL 2.0 with the OpenSSF holding the keys.</p><h3>Where they actually diverge</h3><p>Three years in, this is no longer a rebadged Vault. OpenBao hit v2.5.5 in June and its working groups are building namespaces - a feature Vault locks behind Enterprise - in the open. The flip side cuts just as hard: OpenBao legally can't copy a single line of post-fork Vault code. Every new Vault feature needs a clean-room reimplementation, and that gap compounds every quarter.</p><p>I covered which secrets tool fits which threat model two issues ago - <a href="https://podostack.com/p/issue-025-secrets-in-gitops">issue #025</a> - and everything there applies to both, because the API surface is still compatible. This one's not about threat models. It's about who you trust to still be licensing you the thing in five years.</p><h3>My read</h3><p>Internal use, want vendor support: Vault is fine, BUSL permits you. Building a platform where secrets management is part of the product: Vault is a legal risk you don't need. Betting on Enterprise features arriving free: that's OpenBao's whole reason to exist. Watch whether IBM keeps funding both sides - that's the tell.</p><h3>Links</h3><ul><li><p><a href="https://openbao.org/">OpenBao</a></p></li><li><p><a href="https://github.com/openbao/openbao">GitHub: openbao/openbao</a></p></li><li><p><a href="https://podostack.com/p/issue-025-secrets-in-gitops">Issue #025: secrets in GitOps</a></p></li></ul><div><hr></div><h2>&#128110; The Policy: break-glass access</h2><h3>Why this matters</h3><p>There's a paradox buried in every mature IAM setup: the stronger your SSO, MFA and conditional access get, the more total the lockout when the identity provider goes down. Your admins can't log in to fix the outage because the outage is the thing that checks logins.</p><p>Break-glass accounts are the answer, and most orgs either don't have them or have them configured in a way that fails exactly when needed.</p><h3>The rules that matter</h3><p>Two or more accounts, cloud-only - never synced from on-prem AD, because a dead domain controller must not take your cloud access with it. Explicitly excluded from every conditional access policy. Authentication on hardware FIDO2 keys stored in physical safes, and deliberately a different method than your daily admin accounts use - if Authenticator push is broken for everyone, it's broken for you too.</p><p>Then the rule that turns this from a backdoor into a control: <strong>zero usage</strong>. In normal operation, login count for these accounts is exactly zero, and any successful sign-in fires a critical alert to the whole security team. Not an email. A page.</p><p>Microsoft's guidance on this got refreshed just last month - emergency accounts must be passwordless now - and it includes ready-made alert queries. Steal them. And run a fire drill every 90 days: open the safe, log in, verify the alert actually fired. An untested break-glass account is a password in an envelope and a prayer.</p><h3>Links</h3><ul><li><p><a href="https://learn.microsoft.com/en-us/entra/identity/role-based-access-control/security-emergency-access">Microsoft Entra: emergency access accounts</a></p></li></ul><div><hr></div><p>That's the mix. If one of these four earns a full deep-dive later this year, it'll probably be Timoni - unless the IBM subplot delivers first.</p><p>Questions? Feedback? Reply to this email. I actually read them.</p><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[Postgres statistics: why the planner picks the wrong index]]></title><description><![CDATA[pg_statistic, ANALYZE, default_statistics_target, extended statistics, autovacuum]]></description><link>https://podostack.com/p/postgres-planner-statistics-analyze</link><guid isPermaLink="false">https://podostack.com/p/postgres-planner-statistics-analyze</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Fri, 10 Jul 2026 14:00:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2nIU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2nIU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!2nIU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!2nIU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!2nIU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2nIU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2nIU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!2nIU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!2nIU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!2nIU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49005da8-70f9-4268-8483-152dc88fb29b_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The query had been fast for eight months. A lookup that joined orders to customers, filtered by city and zip, ran in under 10 ms every time, never showed up in <code>pg_stat_statements</code>, nobody thought about it. Then one Tuesday morning the same query started taking nine seconds, and the API behind it started timing out, and the on-call engineer who pulled the plan found a nested loop where there used to be a hash join. Nothing had been deployed. No index had been dropped. What had happened overnight was a 40-million-row bulk import into the orders table, run by a batch job, and that import had quietly invalidated the one thing the planner relies on to make good decisions: its statistics. The query didn't get slower. The planner got dumber, because the numbers it was reasoning from were eight months stale and now described a table that no longer existed.</p><h2>What the planner is actually reading</h2><p>Postgres doesn't plan queries by looking at your data. It plans them by looking at a statistical summary of your data, collected by <code>ANALYZE</code> and stored in the <code>pg_statistic</code> catalog. You'll mostly read it through the friendlier <code>pg_stats</code> view, which decodes the internal arrays into something legible. Everything the planner believes about how many rows a filter will return, whether to use an index, which join algorithm to pick - all of it traces back to four pieces of summary kept per column.</p><p>The first is <code>n_distinct</code>, an estimate of how many distinct values the column holds. A negative value is a ratio - <code>-1</code> means every row is unique, <code>-0.5</code> means roughly half the rows are distinct - which lets the estimate scale as the table grows instead of going stale at a fixed count. The second is the most common values list, <code>most_common_vals</code>, paired with <code>most_common_freqs</code>: the handful of values that appear most often and how frequently each one shows up. This is what lets the planner know that <code>status = 'active'</code> matches 80% of rows while <code>status = 'pending_deletion'</code> matches 0.1%, and treat the two filters completely differently. The third is <code>histogram_bounds</code>, a set of boundary values that divide the rest of the column - everything not already captured in the MCV list - into buckets of roughly equal population, which is how range predicates like <code>created_at &gt; $1</code> get estimated. The fourth is <code>correlation</code>, a number between -1 and 1 describing how closely the physical order of rows on disk matches the sorted order of the column's values.</p><p>That last one quietly decides a lot of index plans, so it's worth dwelling on.</p><h2>Correlation and the cost of an index scan</h2><p>Imagine a table where rows were inserted in <code>created_at</code> order and never moved. The on-disk layout almost exactly matches the sort order of <code>created_at</code>, so correlation is near 1.0. When the planner considers an index scan on a range of dates, it knows the matching heap rows are clustered together on a few adjacent pages, so reading them is nearly sequential and cheap. Now imagine a UUID primary key, where the physical position of a row has nothing to do with the value. Correlation is near zero. An index range scan there means the matching rows are scattered across the whole table, one per page in the worst case, and the planner correctly costs that as expensive random I/O - sometimes expensive enough that it picks a sequential scan of the entire table instead, because reading everything in order beats chasing scattered pages.</p><p>This is why the same index can be a brilliant choice on one column and one the planner refuses to touch on another, and why a freshly <code>CLUSTER</code>-ed table suddenly gets fast index scans it didn't have before. The correlation figure is doing the talking. When it's wrong - because the stats are stale and the table has been reorganized since the last <code>ANALYZE</code> - the planner costs index scans against a layout that no longer exists.</p><h2>How statistics go stale, and why bulk loads are the worst case</h2><p><code>ANALYZE</code> is what refreshes all of this, and in normal operation you never run it by hand because autovacuum runs it for you. The autovacuum daemon watches how many rows have changed in each table and triggers an analyze when the churn crosses a threshold. The formula is <code>autovacuum_analyze_threshold + autovacuum_analyze_scale_factor * reltuples</code> - a small fixed floor (50 by default) plus a fraction of the table size. The scale factor defaults to 0.1, so on a steady table you get a fresh analyze after about 10% of the rows have been inserted, updated, or deleted. For most workloads this is invisible and fine.</p><p>The bulk load is where it breaks, and it breaks in two distinct ways. First, timing: the 40-million-row import dumps all those rows in minutes, but autovacuum doesn't fire mid-transaction, and even after commit there's a delay before the daemon notices and schedules the analyze. In that window - which can be minutes to an hour depending on your autovacuum settings and how busy the daemon is - the planner is working from statistics that describe the table as it was before the import. It thinks the table has the row count it had yesterday, thinks the MCV list still applies, thinks the histogram still covers the real range. Every estimate it makes is anchored to a table that's now a fraction of its current size, and the plans it produces are built on those wrong numbers.</p><p>Second, and nastier: if the bulk load is into a brand-new table, or you've just run a migration that rebuilt one, there may be no statistics at all. With nothing in <code>pg_statistic</code>, the planner falls back to hardcoded guesses baked into the source - a fixed assumption that an equality filter matches a small constant fraction of rows, and so on. Those defaults are occasionally close and usually not, and on a large table "usually not" is a sequential scan where an index would have returned in milliseconds.</p><h2>When one bad estimate cascades</h2><p>The mechanism that turns one bad number into a hundredfold slowdown is worth tracing, because it's never just one bad estimate sitting there harmlessly. The planner builds a plan bottom-up. It estimates how many rows the filter on orders returns, feeds that estimate into the join above it, and the join algorithm it chooses depends entirely on that row count. If it expects the filter to return 12 rows, a nested loop is the obvious choice - for each of 12 rows, probe the index on customers, done in a flash. If it expects 12 rows and the filter actually returns 800,000 because the stats were stale and the value is far more common than the MCV list claimed, that nested loop now runs 800,000 index probes instead of 12. The plan that was optimal for the estimate is catastrophic for the reality. A hash join, which the planner would have chosen if it had known the true row count, would have read each side once and finished in a fraction of the time.</p><p>That's the whole story of the nine-second query. The bulk import made a value far more common than the stale MCV list said it was, the planner underestimated the filter's output by orders of magnitude, picked a nested loop on that underestimate, and the nested loop did hundreds of thousands of times the work it was costed for. The fix was one command - <code>ANALYZE orders</code> - and the query dropped back under 10 ms because the planner could finally see the table it was actually querying. We'd spent twenty minutes reading the query, the index, and the schema, and all three were fine the whole time. What had drifted was the gap between what the planner believed and what was true.</p><h2>The city/zip trap, and extended statistics</h2><p>One class of misestimate survives any amount of fresh single-column statistics, and it's the one that bit our opening query. Postgres, by default, assumes columns are statistically independent. When you filter on two columns at once - <code>WHERE city = 'Springfield' AND zip = '62701'</code> - it estimates the combined selectivity by multiplying the two individual selectivities together. If 0.5% of rows are in Springfield and 0.2% have that zip, it concludes that 0.001% of rows match both, multiplying the fractions as if the two facts were unrelated.</p><p>But city and zip aren't independent at all. A zip code essentially determines its city - knowing the zip tells you the city for free. The rows matching <code>zip = '62701'</code> are almost entirely a subset of the rows matching <code>city = 'Springfield'</code>, so the true combined selectivity is roughly just the zip's selectivity, 0.2%, not the 0.001% the planner computed. The planner underestimates the match count by a couple of hundredfold, and that underestimate feeds straight into the cascade above - it picks a nested loop for what it thinks is a tiny result set, and the result set is two hundred times bigger.</p><p><code>CREATE STATISTICS</code> is the fix, available since Postgres 10. You declare an extended statistics object over the correlated columns and let <code>ANALYZE</code> measure their actual relationship:</p><pre><code>CREATE STATISTICS orders_city_zip (dependencies, ndistinct)
  ON city, zip FROM orders;
ANALYZE orders;</code></pre><p>The <code>dependencies</code> kind teaches the planner the functional dependency between city and zip, so it stops multiplying their selectivities as if they were independent. The <code>ndistinct</code> kind tracks the number of distinct combinations across the column group, which fixes underestimates in <code>GROUP BY</code> over multiple columns. These extended objects are computed from the same row sample <code>ANALYZE</code> already takes, so they cost almost nothing extra to maintain - but they don't exist until you create them, and Postgres will never create them for you. This is the single most common missing-statistics object in production, and the symptom is always the same: a multi-column filter on related columns that the planner wildly underestimates.</p><h2>default_statistics_target and when to raise it</h2><p>The resolution of all this is governed by <code>default_statistics_target</code>, which defaults to 100. That number controls how many entries go in the MCV list and how many buckets the histogram gets - effectively, how detailed a picture <code>ANALYZE</code> paints. For most columns 100 is plenty. For a column with a long, skewed tail of values - where the top 100 most common values don't capture enough of the distribution, and the planner keeps misjudging the less-common ones - raising it helps the estimates a lot.</p><p>You can raise it globally, but that's usually the wrong move; it bloats <code>pg_statistic</code> and slows <code>ANALYZE</code> across every column whether they need it or not. The targeted lever is per-column: <code>ALTER TABLE orders ALTER COLUMN customer_id SET STATISTICS 1000</code>. Now that one column gets a 1000-entry MCV list and a finer histogram on the next analyze, and nothing else pays for it. Reach for this when <code>EXPLAIN ANALYZE</code> shows a large gap between estimated and actual rows on a filter against a high-cardinality, skewed column, and the stats are confirmed fresh - raising the target on a column that's simply stale just produces a more detailed wrong answer.</p><h2>Why ANALYZE is cheap and belongs in every migration</h2><p>The thing that makes the stale-stats failure so frustrating is how trivial the fix is. <code>ANALYZE</code> reads a sample of the table - not the whole thing, a statistically sufficient sample sized from the statistics target - computes the summaries, and writes them to <code>pg_statistic</code>. On a large table it's seconds, not the minutes a full <code>VACUUM FULL</code> or reindex would take. Crucially, it takes only a light lock. <code>ANALYZE</code> acquires a <code>SHARE UPDATE EXCLUSIVE</code> lock, which lets reads and writes continue against the table the whole time - it does not take the exclusive lock that blocks traffic, so there's no reason to fear running it on a live system during business hours.</p><p>Given that, the rule writes itself: any operation that meaningfully changes a table's contents should be followed by an explicit <code>ANALYZE</code>. A bulk import, a large backfill, a migration that rewrites a column, a <code>pg_restore</code> into a fresh database - run <code>ANALYZE</code> on the affected tables before you let production traffic hit them, rather than waiting the minutes-to-an-hour for autovacuum to notice. It costs seconds and it closes exactly the window where the planner is reasoning from a table that no longer exists. The opening incident was nine seconds of timeout per request for the better part of an hour because the batch job didn't end with one cheap command.</p><h2>The decision framework</h2><p>When a query that was fine goes sideways, the question is which layer lied, and the order you check matters.</p><p>Pull the plan first with <code>EXPLAIN ANALYZE</code> and compare estimated rows against actual rows at each node. A close match all the way up means the planner saw the table correctly and the plan is genuinely the best available - the problem is elsewhere, maybe a missing index or the query itself. A large gap at some node is the planner working from bad numbers, and that gap is where you focus.</p><p>If the gap is on a single-column filter, suspect staleness first. Check <code>last_analyze</code> and <code>last_autoanalyze</code> in <code>pg_stat_user_tables</code>, and if they predate a recent bulk change, run <code>ANALYZE</code> and re-check the plan before doing anything more elaborate. If the column is high-cardinality and skewed and the stats are fresh, that's the case for raising its per-column statistics target. If the gap is on a filter spanning two or more related columns - the city/zip shape - no single-column tuning will touch it, and the answer is a <code>CREATE STATISTICS</code> object over the group. If the gap is specifically on index-versus-sequential choice and the table was recently reorganized, look at the correlation figure, which goes stale the same way everything else does.</p><h2>The ones that page me at 3 AM</h2><p>Roughly in order of how often each has cost me a night:</p><ul><li><p>Bulk loading and walking away. The import commits, autovacuum hasn't fired yet, and for the next stretch every plan is built on pre-import statistics. End the load with an explicit <code>ANALYZE</code> and the window never opens.</p></li><li><p>Then the independence assumption, where it doesn't hold. City and zip, country and currency, product and category - any pair where one implies the other will be underestimated, sometimes by hundreds of times, until a <code>CREATE STATISTICS</code> object teaches the planner the dependency.</p></li><li><p>Raising <code>default_statistics_target</code> globally to fix one column slows every analyze and bloats the catalog for no benefit on the columns that were already fine. Set it per-column where <code>EXPLAIN ANALYZE</code> proves the need.</p></li><li><p>Treating a nested-loop blowup as a query problem misses the point. The nested loop wasn't wrong for the row count the planner estimated - the estimate was. Fix the cardinality and the join algorithm fixes itself.</p></li><li><p>Fear that <code>ANALYZE</code> will lock the table keeps people from running it, but it takes a <code>SHARE UPDATE EXCLUSIVE</code> lock, reads and writes continue throughout, and it runs in seconds. No reason to schedule it for a maintenance window.</p></li><li><p>A restored or freshly-migrated database has no statistics at all, which people forget. Until the first <code>ANALYZE</code>, the planner runs on hardcoded default guesses, which on a large table means sequential scans where indexes existed all along.</p></li><li><p>And reading a misestimate without checking <code>last_analyze</code>. Half the "the planner is broken" reports are simply stale statistics, and the check is one query against <code>pg_stat_user_tables</code> before you go hunting for anything subtler.</p></li></ul><p>The planner is only as good as its picture of the data, and that picture is a snapshot that ages. Most of the dramatic, overnight, nothing-changed query regressions you'll chase come down to the snapshot having drifted away from reality - a bulk load that outran autovacuum, a correlated pair the planner was multiplying blind, a restored database nobody analyzed. The catalog tables tell you which one it is, and <code>ANALYZE</code> is almost always the cheapest fix you'll apply all week.</p>]]></content:encoded></item><item><title><![CDATA[Redis memory: why 16 GB of data needs 32 GB of RAM]]></title><description><![CDATA[Fragmentation, RSS, jemalloc, eviction, fork overhead]]></description><link>https://podostack.com/p/redis-memory-fragmentation</link><guid isPermaLink="false">https://podostack.com/p/redis-memory-fragmentation</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Wed, 08 Jul 2026 14:01:14 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!j9pU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!j9pU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!j9pU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!j9pU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!j9pU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!j9pU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!j9pU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!j9pU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!j9pU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!j9pU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6da778d9-d1fe-4b88-8317-38c4259cfb1e_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The box had 32 GB. The dashboard said Redis was holding 16 GB of data, comfortably under half. Then the OOM killer took it at 03:00, and the on-call engineer spent the first twenty minutes convinced the graph was lying, because how does a process reported at 16 GB get killed on a 32 GB machine. The graph wasn't lying. It was just answering a different question than the one being asked. <code>used_memory</code> was 16 GB. The resident set the kernel actually accounted for was sitting near 31 GB, and Redis had no idea, because the bytes it gave back to its allocator never made it back to the OS. This is the single most expensive misunderstanding about running Redis at scale, and it applies one-for-one to Valkey, which forked from the same codebase and inherited the same allocator behavior.</p><h2>Two numbers that should be one</h2><p>Redis reports memory through <code>INFO memory</code>, and the first thing to internalize is that the two headline numbers measure different layers of the stack. <code>used_memory</code> is what the allocator handed Redis for its data and bookkeeping - the logical footprint. <code>used_memory_rss</code> is the resident set size the operating system sees, the physical pages actually mapped to the process. In a perfect world they'd track each other. They don't, and the gap is where the trouble lives.</p><p>The ratio between them, <code>mem_fragmentation_ratio</code>, is <code>used_memory_rss</code> divided by <code>used_memory</code>. People read it as a fragmentation gauge and mostly that's right, but it's quietly overloaded. A value around 1.0 to 1.1 is healthy. Up around 1.5 means roughly half again as much physical memory as your data needs, which is what bit the box above. The part that catches teams off guard is the other direction. A ratio below 1.0 is not "great, negative fragmentation." It means part of Redis has been pushed to swap, so the OS-resident portion is now smaller than the logical data, and the difference is sitting on disk. That's the worst state on this whole page. A swapping Redis has tail latencies that look like the network is on fire, and a fragmented Redis just wastes RAM. I'll take the wasted RAM.</p><p>One more number worth pulling: <code>mem_fragmentation_bytes</code>, the absolute gap. On a small instance a ratio of 1.4 is a rounding error in bytes; on a 60 GB instance it's a second machine's worth of memory you're paying for and not using.</p><h2>Why the allocator keeps what you delete</h2><p>The reason RSS doesn't fall when you delete keys comes down to jemalloc, the allocator Redis and Valkey ship with by default. jemalloc doesn't hand out memory in the exact size you ask for. It rounds every request up into one of a fixed set of size classes - 8, 16, 32, 48, 64 bytes and so on up through larger spacings. Ask for a 33-byte value and you get a 48-byte slot, and those 15 bytes are gone to internal fragmentation before any key has been deleted. Across millions of values of awkward sizes, that overhead alone is real money.</p><p>The deletion problem is worse and less intuitive. jemalloc groups allocations of the same size class onto runs of contiguous pages. When you delete a key, its slot goes back onto jemalloc's free list for that size class - available for the next value of that size, but still resident, still counted in RSS. The OS only gets a page back when an entire run of pages becomes free and jemalloc decides to return it. If one live 64-byte object remains on a page full of freed 64-byte slots, that whole page stays mapped. So a workload that fills to 30 GB, deletes 90% of its keys, and sits there will keep reporting an RSS near 30 GB while <code>used_memory</code> drops to 3 GB. The data is gone. The memory isn't. This is exactly why restarting Redis "fixes" a memory problem - the restart rebuilds the heap from scratch with no holes in it, and for a while the ratio is beautiful again. It's not a fix, it's a reset, and the fragmentation comes right back as the access pattern resumes.</p><h2>Active defragmentation, and what it costs</h2><p>Redis 4.0 added <code>activedefrag</code> precisely for the stuck-RSS case. With jemalloc compiled in, Redis can ask the allocator which objects are sitting on sparsely-used pages and then copy those live objects to freshly allocated, densely-packed memory, freeing the old pages back to the OS. It runs incrementally in a background cycle so it doesn't stall the main thread, governed by thresholds like <code>active-defrag-ignore-bytes</code> and <code>active-defrag-threshold-lower</code>, and a CPU ceiling so it backs off under load.</p><p>It works, but it isn't free and it isn't instant. Defrag burns CPU walking and relocating objects, and on a busy instance you'll watch the fragmentation ratio drift down over minutes to hours rather than snap. I've seen people enable it, check the ratio thirty seconds later, see no change, and conclude it's broken. Give it time. There's also a manual lever, <code>MEMORY PURGE</code>, which tells jemalloc to release as much free-but-resident memory as it can right now - useful as a one-shot after a big mass deletion, but it doesn't relocate live objects the way active defrag does, so it can't unstick a page that still holds one survivor.</p><h2>The fork that doubles you</h2><p>Here's the operational landmine that turns a comfortable instance into an OOM kill. Redis persists with RDB snapshots and AOF rewrites, and both work by calling <code>fork()</code>. The child gets a copy-on-write view of the parent's memory: pages are shared until one side writes to them, at which point the kernel copies that page so the two processes diverge. If your instance is idle during the save, COW costs almost nothing. If your instance is taking writes during the save - and a production cache always is - every page the parent touches gets duplicated, and in the pathological case the child ends up holding a near-complete second copy of the dataset.</p><p>That's the real shape of "16 GB needs 32 GB." It isn't just fragmentation. A 16 GB dataset under steady write load, caught mid-RDB-save, can briefly push total RSS toward 32 GB between parent and child. If you sized the box for the dataset plus a fragmentation margin and forgot the fork, the save itself is what tips you into the killer. This is why <code>vm.overcommit_memory = 1</code> is in every Redis production guide - without it, the <code>fork()</code> can be refused outright when the kernel does naive accounting, and your background save fails. The deeper fix is to leave real headroom: plan for the dataset, plus fragmentation, plus a fork copy under your peak write rate.</p><h2>maxmemory, eviction, and the OOM that comes back as an error</h2><p>If you don't set <code>maxmemory</code>, Redis grows until the kernel kills it. So you set it. But <code>maxmemory</code> plus the wrong policy trades one outage for another. The default policy is <code>noeviction</code>: at the limit, reads still work, but any write that would grow memory gets rejected with an OOM error returned to the client. For a cache fronting a database, that's usually catastrophic - the app starts taking write errors under exactly the load spike that filled the cache, and a "caching layer" becomes a hard dependency that fails closed.</p><p>The eviction policies are the alternative, and which one fits depends on what the instance is for. <code>allkeys-lru</code> and the newer <code>allkeys-lfu</code> evict across the whole keyspace by approximate recency or frequency - the right default for a pure cache where everything is disposable. The <code>volatile-*</code> family (<code>volatile-lru</code>, <code>volatile-lfu</code>, <code>volatile-ttl</code>, <code>volatile-random</code>) only evicts keys that carry a TTL, which is what you want when the instance mixes throwaway cache entries with keys that must persist. The trap in the volatile policies is that if nothing currently has a TTL, there's nothing eligible to evict, and you're back to OOM errors despite having "configured eviction." LRU and LFU here are approximations, by the way - Redis samples a handful of keys rather than maintaining a true global ordering, because exact LRU across millions of keys would cost more memory than it saves.</p><p>Eviction and fragmentation interact in a way that surprises people. <code>maxmemory</code> is enforced against <code>used_memory</code>, the logical number, not against RSS. So an instance can be evicting hard - throwing away useful data to stay under the limit - while RSS sits well above <code>maxmemory</code> because of fragmentation holes the allocator won't release. You're losing cache hits and over-consuming RAM at the same time. When you see eviction counts climbing on an instance that the dashboard says has spare physical memory, fragmentation is usually the reason the two views disagree.</p><h2>Big keys and why you never run KEYS</h2><p>A single hash, set, or sorted set with millions of members is its own memory hazard. It allocates as one logical structure, it can't be partially evicted - eviction works at the key granularity, so a 4 GB sorted set is all-or-nothing - and deleting it synchronously blocks the main thread while Redis frees every member, which is why <code>UNLINK</code> (background free) exists alongside <code>DEL</code>. Big keys also wreck the COW math during forks, because one write into a giant value can dirty a lot of pages at once.</p><p>Finding them is where the second classic outage hides. The instinct is <code>KEYS *</code> to list the keyspace, and <code>KEYS</code> is O(n) and blocks the single main thread for the entire scan. On a large instance that's a multi-second stall where Redis answers nothing, and you've created an incident while investigating one. <code>SCAN</code> is the answer - a cursor-based iterator that returns small batches and lets other commands interleave. Same rule for the type-specific variants, <code>HSCAN</code>, <code>SSCAN</code>, <code>ZSCAN</code>, when you're digging into one big collection. For sizing, <code>redis-cli --bigkeys</code> and <code>--memkeys</code> walk the keyspace with SCAN under the hood and report the largest keys per type without freezing anything.</p><h2>The decision framework</h2><p>The numbers tell you which problem you have. Read them in this order before you touch a config.</p><p>Start with the fragmentation ratio. If it's below 1.0, stop everything else - you're swapping, and the fix is more RAM or a smaller dataset, not defrag. If it's comfortably between 1.0 and 1.3 and RSS fits the box with fork headroom, you don't have a memory problem and you should resist the urge to tune. If it's elevated, 1.4 and up, and the absolute <code>mem_fragmentation_bytes</code> is large enough to matter, that's when <code>activedefrag yes</code> earns its CPU.</p><p>Size the box for three things stacked, not one. The dataset is the floor. Add a fragmentation margin on top - 20 to 50% depending on how churn-heavy and awkwardly-sized your values are. Then add room for a fork copy proportional to your write rate during saves. A 16 GB dataset on a 16 GB box is a future incident with a date on it.</p><p>Pick the eviction policy from what the instance actually is. Pure disposable cache: <code>allkeys-lru</code> or <code>allkeys-lfu</code>, and never <code>noeviction</code>. Mixed cache-and-state where some keys must survive: a <code>volatile-*</code> policy, but only if you're disciplined about setting TTLs, because an empty eligible set lands you back on OOM errors. Durable store you don't want silently shrinking: <code>noeviction</code> on purpose, with hard alerting on <code>used_memory</code> so a human intervenes before the writes start failing.</p><h2>Common mistakes</h2><p>The same handful of errors show up across nearly every Redis memory incident.</p><ul><li><p>Sizing the instance for <code>used_memory</code> and ignoring RSS. The kernel kills on RSS, and RSS carries fragmentation plus any fork copy. The logical number is the smallest of the three.</p></li><li><p>Reading a sub-1.0 fragmentation ratio as good news. It means swap, which is the worst outcome on this page, and it disguises itself as a low number.</p></li><li><p>Running <code>noeviction</code> on a cache. The first traffic spike that fills it turns the cache into a write-rejecting hard dependency, and the database behind it loses the buffer right when it needs it most.</p></li><li><p>Setting a <code>volatile-*</code> policy and then not setting TTLs. Nothing is eligible to evict, so the instance OOMs anyway and the eviction config reads like a safety net that was never connected.</p></li><li><p>Forgetting the fork. The instance runs fine for weeks, then OOM-kills exactly during an RDB save or AOF rewrite under write load, and the postmortem blames "a memory leak" instead of copy-on-write.</p></li><li><p>Running <code>KEYS *</code> to find big keys. O(n) on the main thread, a multi-second stall, an outage created by the investigation. <code>SCAN</code> and <code>--bigkeys</code> exist for this.</p></li><li><p>Restarting to "fix" memory and calling it done. The restart resets the heap and the ratio looks great, but the same access pattern rebuilds the same fragmentation, and you're back in a week without having changed anything.</p></li><li><p>Enabling <code>activedefrag</code> and judging it in thirty seconds. It works over minutes to hours and costs CPU while it runs. Watch the ratio trend, not a single sample.</p></li></ul><p>None of this changes whether you're on Redis or Valkey - same allocator, same fork, same eviction model, same numbers in <code>INFO</code>. The mental model that saves you is small and unglamorous: there are three memory figures, not one, the OS only cares about the largest, and the gap between logical and resident is paid for in real RAM whether or not anything is using it.</p>]]></content:encoded></item><item><title><![CDATA[Issue #025 - Four ways to put secrets in Git: three are wrong for you]]></title><description><![CDATA[SOPS and age, Sealed Secrets, External Secrets Operator, Vault and OpenBao, threat models]]></description><link>https://podostack.com/p/issue-025-secrets-in-gitops</link><guid isPermaLink="false">https://podostack.com/p/issue-025-secrets-in-gitops</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Tue, 07 Jul 2026 14:02:05 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/dab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!VviJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!VviJ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!VviJ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!VviJ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!VviJ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!VviJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!VviJ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!VviJ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!VviJ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!VviJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdab3eb8c-076f-47cd-a594-6dfd29f94736_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The repo had been public-inside-the-company for years, the kind every platform team has - charts, kustomizations, the lot. I was greping it for an unrelated <code>ConfigMap</code> and the match landed in a file called secret.yaml, committed in 2021, with a data block that was very obviously base64. I ran it through base64 -d out of reflex, the way you do, expecting nothing. It was the production database password for a service that was still very much running. Not a sealed anything, not encrypted - a plain Kubernetes Secret, which is to say base64, which is to say plaintext with an extra step, sitting in Git history forever, readable by everyone who'd ever had a clone.</p><p>The uncomfortable part wasn't that someone did it. The uncomfortable part was that for three years the repo's whole story had been "secrets live in Git, that's our GitOps model, it's fine." Nobody had lied. They'd just never written down what "fine" was protecting against, so base64 had quietly counted as protection. The fix that afternoon was easy. The thing that took longer was the question underneath: if we're putting secrets in Git on purpose - and GitOps means we are - what exactly are we defending against, and which of the four standard tools actually defends against that?</p><p>Because there are four. SOPS, Sealed Secrets, External Secrets Operator, HashiCorp Vault and its operators. Every team picks one, usually the one a teammate used at their last job, and then defends it as if it were the obviously-correct choice. It isn't. Three of the four are wrong for any given team, and which three depends entirely on a threat model most teams never write down.</p><h2>&#127959;&#65039; Architectural Pattern: where the trust root lives and what's in Git</h2><h3>The question nobody writes down</h3><p>Every one of these four tools answers the same question - "how do I get a secret from my laptop to a running pod, through a Git repo, without it leaking" - and they answer it completely differently because they disagree about one thing: where the key that can read the secret lives, and therefore who, exactly, you're trusting.</p><p>When I finally drew this on a whiteboard for my own team, the question that mattered wasn't "which is most secure" - that one has no answer. It was two narrower ones: <strong>what is actually committed to Git, and what can turn it back into a plaintext secret.</strong> Once I had those pinned, the right tool stopped being a debate. The 2021 <code>secret.yaml</code> was what you got by skipping them - a model nobody chose, defending against nothing in particular.</p><p>The four sort into two camps along that axis, and in practice the camp mattered more than the specific tool.</p><h3>Camp one: the ciphertext is in Git</h3><p>SOPS and Sealed Secrets both put encrypted bytes into the repo. The secret material is there, in a commit, forever. What differs is who holds the key that reverses it.</p><p>With <strong>SOPS</strong>, you encrypt each value against a key you control - an age keypair, or a cloud KMS key (AWS KMS, GCP KMS, Azure Key Vault), or PGP. The repo holds ciphertext; the decryption key lives wherever you put it. The threat this defends against is precise and limited: someone reading your Git repo learns nothing, because the values are encrypted. But anyone who holds the age private key or can call that KMS key decrypts everything, including every secret that ever appeared in history. The trust root is the key, and the key is off in KMS or on a few laptops. Recover the key, recover the entire history.</p><p>With <strong>Sealed Secrets</strong>, the trust root moves into the cluster. The Bitnami controller generates an RSA keypair on first boot and keeps the private key as a Secret in <code>kube-system</code>. You encrypt with the public half using <code>kubeseal</code> - anyone can, the public key is public - and only that one controller, holding that one private key, can ever decrypt. The ciphertext in Git is genuinely useless to anyone who isn't that controller. That's a stronger property than SOPS in one specific way: there's no decryptable-by-a-human key sitting in KMS. And it's a much sharper footgun in another, which I'll get to, because that private key is now the single most important thing in your cluster and most teams don't know where it's backed up.</p><h3>Camp two: the secret is not in Git at all</h3><p>The other two refuse the premise. Nothing secret goes into the repo - only a reference.</p><p><strong>External Secrets Operator</strong> commits an <code>ExternalSecret</code> manifest that says, in effect, "the value for <code>db-password</code> lives in AWS Secrets Manager at this path, go fetch it." ESO runs in-cluster, authenticates to the external store, pulls the real value, and materializes a normal Kubernetes Secret. Git holds a pointer and nothing else. Clone the repo, get the entire repo, and you have learned the <em>names</em> of secrets and where they're stored - never a byte of the secrets themselves. The trust root is the external store plus ESO's credential to read it.</p><p><strong>Vault</strong> (via the Vault Secrets Operator, the CSI provider, or Vault Agent) is the same shape - reference in Git, value in the store - with one capability the other three structurally can't have: the secret can be <em>dynamic</em>. Vault can mint a Postgres credential that exists for one hour and then revokes itself. There's no long-lived password to leak, because by the time anyone exfiltrates it, it's dead. That's a different threat model again: not "protect the secret in transit through Git" but "minimize the lifetime of the thing worth stealing."</p><p>So the four are really two questions stacked. First: is the ciphertext in Git (SOPS, Sealed Secrets) or is only a reference in Git (ESO, Vault)? Second, within each camp: who holds the key, or how long does the secret live? Answer those honestly about your own threat model and the field of four collapses to one. Answer them with "it's what my last team used" and you're back to base64.</p><h3>Links</h3><ul><li><p><a href="https://github.com/getsops/sops">SOPS on GitHub (getsops)</a></p></li><li><p><a href="https://github.com/bitnami-labs/sealed-secrets">Sealed Secrets on GitHub</a></p></li><li><p><a href="https://external-secrets.io/latest/">External Secrets Operator</a></p></li><li><p><a href="https://developer.hashicorp.com/vault/docs/deploy/kubernetes/vso">Vault Secrets Operator docs</a></p></li></ul><h2>&#127386; The Showdown: SOPS vs Sealed Secrets vs ESO vs Vault</h2><p>I've run three of these four in anger and watched a team get badly burned by the fourth, so let me be specific about where each one earns its place and where it bites.</p><h3>SOPS plus age or KMS</h3><p>SOPS is the one I reach for when a team wants encryption-in-Git and doesn't already run anything heavier. It's a CNCF Sandbox project now - donated out of Mozilla in 2023, after seven-odd years of life, picked up by a fresh group of maintainers under the <code>getsops</code> org. That matters because the Mozilla-era SOPS had gone quiet and a lot of people wrote it off; it's actively maintained again.</p><p>The thing SOPS gets right that nothing else does: it encrypts <em>values</em> and leaves <em>keys</em> in plaintext. Your <code>secret.yaml</code> still diffs cleanly in a PR - you can see that <code>db-password</code> changed without seeing what it changed to. No in-cluster component is needed for the encryption itself, which is why Flux integrates it natively: Flux decrypts at reconcile time with a key you've handed the controller. ArgoCD is clumsier - you bolt on <code>ksops</code> as a Kustomize plugin or run a SOPS-aware sidecar, and it's never as clean as the Flux path.</p><p>Where it bites: the decryption key is a human-holdable thing. An age private key on a laptop, a KMS key a CI role can call. Whoever has it reads the whole repo's history. SOPS gives you encryption-at-rest in Git and not one thing more, and teams routinely forget that and treat "it's SOPS-encrypted" as if it meant "it's safe even if the key leaks." It does not.</p><h3>Sealed Secrets</h3><p>Sealed Secrets is the simplest mental model: encrypt to the cluster, only the cluster decrypts. The controller's RSA private key never leaves <code>kube-system</code>. You seal with <code>kubeseal</code> and the public key, commit the SealedSecret, and the controller unseals it into a real Secret. By default the seal is <em>strict</em> - it binds the ciphertext to the exact namespace and name, so a secret sealed for prod/db-password won't decrypt as staging/db-password, which is a genuinely nice property and the source of most "why won't this decrypt" tickets when someone renames a namespace. There are namespace-wide and cluster-wide scopes if you need to relax that, and relaxing it is a deliberate, auditable choice.</p><p>Here's the part that bit a team I watched. They rebuilt a cluster - new nodes, restore the GitOps repo, let it reconcile. Every <code>SealedSecret</code> came back as a decryption error and nothing started. The controller in the new cluster had generated a <em>new</em> keypair, and the ciphertext in Git was sealed against the <em>old</em> private key, which had only ever existed inside the cluster they'd just torn down. Nobody had backed it up, because nobody had internalized that the controller's private key was now the single most precious object they owned. They got lucky - the old cluster was still half-alive and they pulled the key out. If it hadn't been, every production secret would have had to be regenerated by hand. Sealed Secrets moves the trust root somewhere very safe and then dares you to forget where you put the only copy.</p><p>One bit of good news on maintenance, because there was a scare. When Broadcom restructured the Bitnami catalog on August 28, 2025 - the same change that broke a wave of <code>bitnami/*</code> Helm charts and container images industry-wide, the one I mentioned back in the Valkey issue - a lot of people assumed Sealed Secrets was caught in it. It wasn't. Sealed Secrets lives under <code>bitnami-labs</code>, not the commercial Bitnami catalog, its images stay on <code>docker.io/bitnami</code>, and it kept shipping. If you deferred adopting it last autumn out of license nerves, that fear was misplaced.</p><h3>External Secrets Operator</h3><p>ESO is the one I'd call the default for most teams that already have a secret store, and it's also the one with the most interesting 2025. The pitch is clean: nothing secret in Git, ever, just <code>ExternalSecret</code> references, and the operator syncs from AWS Secrets Manager, GCP Secret Manager, Azure Key Vault, Vault, and a long tail of others. Compromising the Git repo gets an attacker the topology of your secrets and not the secrets. It's a CNCF Sandbox project, accepted back in July 2022.</p><p>And in August 2025 it nearly died. On the 13th the maintainers announced they were pausing all releases - no features, no security patches, no image publishing - because the project had effectively <em>one</em> maintainer carrying it quasi-full-time and he'd burned out. The quote that stuck with me was Gergely Brautigam's, roughly: people feel entitled to say "I use this, it's broken, fix it," and that's not how any of this works. For a few months ESO's future was a live question, which is a deeply uncomfortable thing to read about a tool you've wired into every cluster you run. By mid-2026 it recovered - new maintainers stepped up, releases resumed, it's actively maintained again. But the episode is the most honest argument in this whole comparison: the trust root for ESO isn't only the external store, it's also the bus factor of an operator that spent a season at exactly one.</p><h3>Vault and the OpenBao question</h3><p>Vault is the heavyweight, and the only one of the four that does <em>dynamic</em> secrets - credentials minted on demand with a lease, auto-revoked when the lease ends. The Vault Secrets Operator understands Vault's lease model natively: it renews leases, caches the client token, and rotates dynamic database credentials at 67% of their TTL by default, in-cluster, declaratively. If your threat model is "a stolen credential should be worthless within the hour," nothing else here competes, because the other three all hand you a long-lived secret and call it done.</p><p>The cost is operational weight - Vault is a system you run, seal, unseal, back up, and staff, not a controller you <code>helm install</code> and forget. And there's a license asterisk that's now load-bearing. HashiCorp moved Vault to the BUSL on August 10, 2023 (1.15 was the first BUSL release, 1.14 the last under MPL), and IBM closed its acquisition of HashiCorp on February 27, 2025, so Vault is an IBM product under a source-available license today. The open-source answer is <strong>OpenBao</strong>, forked from Vault 1.14.0 - the last MPL release - now under the Linux Foundation and an OpenSSF Sandbox project since June 2025. If you want Vault's model without the BUSL, OpenBao is the fork to look at, the same way Valkey was the fork for Redis. As of mid-2026 the BUSL hasn't been reversed and there's no sign it will be, so the choice is real, not hypothetical.</p><h3>Links</h3><ul><li><p><a href="https://www.cncf.io/projects/sops/">SOPS in CNCF</a></p></li><li><p><a href="https://github.com/cncf/toc/issues/1819">ESO release pause announcement (cncf/toc #1819)</a></p></li><li><p><a href="https://external-secrets.io/latest/introduction/stability-support/">ESO stability and support</a></p></li><li><p><a href="https://thenewstack.io/meet-openbao-an-open-source-fork-of-hashicorp-vault/">Meet OpenBao, the open-source Vault fork</a></p></li><li><p><a href="https://developer.hashicorp.com/vault/tutorials/kubernetes-introduction/vault-secrets-operator">Vault Secrets Operator with dynamic secrets</a></p></li></ul><h2>&#128110; The Policy: a manifest each, and which threat model picks which</h2><p>Enough comparison. Here's what the two camps actually look like on disk, and then the part that matters - which threat model lands you on which.</p><p>A SOPS-with-age secret, after <code>sops --encrypt</code>, keeps the structure readable and hides only the values:</p><pre><code>apiVersion: v1
kind: Secret
metadata:
  name: db-creds
  namespace: prod
data:
  password: ENC[AES256_GCM,data:9f3k...,iv:7a2c...,tag:b1e4...,type:str]
sops:
  age:
    - recipient: age1ql3z7hjy54pw3hyww5ayyfg7zqgvc7w3j2elw8zmrj2kg5sfn9aqmcac8j
      enc: |
        -----BEGIN AGE ENCRYPTED FILE-----
        ...
        -----END AGE ENCRYPTED FILE-----</code></pre><p>That commits to Git as-is. The <code>recipient</code> line is the age public key - safe to publish. The decryptable half lives on the laptops and CI roles you've handed the private key to.</p><p>An ESO reference commits no secret material at all - just a pointer and a refresh interval:</p><pre><code>apiVersion: external-secrets.io/v1
kind: ExternalSecret
metadata:
  name: db-creds
  namespace: prod
spec:
  refreshInterval: 1h
  secretStoreRef:
    name: aws-secrets-manager
    kind: SecretStore
  target:
    name: db-creds
  data:
    - secretKey: password
      remoteRef:
        key: prod/db/password</code></pre><p>Clone that repo and you've learned that a thing called <code>prod/db/password</code> exists in a store called <code>aws-secrets-manager</code>. You have learned nothing you could use.</p><p>When I've had to make this call for real, I started from the threat model rather than the tool I already knew:</p><p><strong>You're air-gapped or have no external secret store, and you want the simplest GitOps model.</strong> Sealed Secrets. The trust root is the in-cluster controller key, you depend on nothing external, and <code>kubeseal</code> is a five-minute setup. Just write down where the controller's private key is backed up <em>before</em> you ship anything, or the cluster-rebuild story from the Showdown becomes yours.</p><p><strong>You're air-gapped or store-less but want readable diffs and multi-recipient encryption</strong> (different keys for prod and staging, say). SOPS with age. Lighter than Sealed Secrets if you're already on Flux, and the encrypt-the-value-keep-the-key property is worth a lot in review. Accept that the key is human-holdable and guard it accordingly.</p><p><strong>You already run a secret store</strong> - AWS Secrets Manager, GCP, Azure, or Vault - and you want nothing secret in Git. ESO. It's the boring-enterprise default for a reason, and the recovery in 2026 means you can stop worrying about it disappearing. Trust shifts to the store and ESO's read credential, which is where most mature teams want it.</p><p><strong>Your threat model is credential lifetime</strong> - the secret worth stealing should be dead within the hour, or you need per-request database credentials for compliance. Vault, via VSO, with dynamic secrets. Or OpenBao if the BUSL is a dealbreaker. This is the only camp that makes the stolen-secret problem go away by construction, and it's also the only one that's a system you have to staff. Don't pick it for static API keys you could've put in any of the other three.</p><p>The trap, every time, is picking by familiarity and then writing the threat model backwards to justify it. Pick the threat model first. Three of these four will be visibly wrong for it, and the fourth will feel obvious.</p><h3>Links</h3><ul><li><p><a href="https://github.com/getsops/sops#22encrypting-using-age">SOPS with age and KMS</a></p></li><li><p><a href="https://github.com/bitnami-labs/sealed-secrets#scopes">Sealed Secrets scopes (strict, namespace-wide, cluster-wide)</a></p></li><li><p><a href="https://external-secrets.io/latest/api/externalsecret/">ESO ExternalSecret API reference</a></p></li><li><p><a href="https://developer.hashicorp.com/vault/docs/secrets/databases">Vault dynamic database secrets</a></p></li></ul><h2>What secrets actually are</h2><p>Strip the tooling away and a secret is a trust root - the thing everything downstream believes without checking. Every one of these four tools is really an argument about where that root should sit: in a KMS key, in a controller's memory, in an external store, in a one-hour lease. They aren't competing on features. They're competing on where you're willing to put the thing that, if it leaks, ends you.</p><p>That base64 in the 2021 repo was a trust root nobody had decided to place anywhere, so it had ended up everywhere - in the Git history, in every clone, in the laptop backups of everyone who ever pulled the repo. Picking SOPS or Sealed Secrets or ESO or Vault is, underneath, the act of deciding. The wrong tool for your threat model still beats the non-decision that base64 represents, by a wide margin.</p><p>Go find one secret in your repos that nobody ever decided where to trust. There's at least one. See you Tuesday.</p><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[Iceberg vs Delta vs Hudi: the table-format war is over]]></title><description><![CDATA[Open table formats, metadata layers, snapshots, REST catalogs, interop]]></description><link>https://podostack.com/p/data-lake-table-formats-iceberg-delta-hudi</link><guid isPermaLink="false">https://podostack.com/p/data-lake-table-formats-iceberg-delta-hudi</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Fri, 03 Jul 2026 14:01:26 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5qwd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5qwd!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!5qwd!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!5qwd!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!5qwd!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5qwd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5qwd!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!5qwd!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!5qwd!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!5qwd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff4552512-d116-4a53-a4fe-33bdbe65bc90_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>A team I talked to picked Hudi in 2022 and spent the early months of 2026 migrating off it. Not because Hudi broke. It ran fine for three years. They moved because every tool they wanted to bring in - a new query engine, a managed warehouse for the analysts, a catalog the security team would sign off on - assumed Iceberg, and they were tired of being the one workload that needed a special path. The migration wasn't a verdict on Hudi's engineering. It was the team noticing, a couple of years late, that the question they'd argued about in 2022 had quietly stopped being the question. Picking a table format used to feel like picking a database - a decision you'd live with for a decade. By 2026 it feels more like picking which of three nearly-identical things to write your data in, because the engines have learned to read all of them anyway.</p><p>That's the short version of where we are: the format war is basically over, Iceberg won the mindshare, and almost nobody should care as much as they did. The longer version is more interesting, because the war didn't end with one format killing the others. It ended with convergence, and the real fight moved somewhere else.</p><h2>What a table format even is</h2><p>Strip away the branding and all three - Iceberg, Delta Lake, Hudi - solve the same problem. You have a pile of Parquet files in object storage. Object storage knows nothing about tables, transactions, or schemas; it knows about objects and prefixes. A table format is the metadata layer you bolt on top so that a directory of Parquet files behaves like a real table - one you can update, query consistently, evolve, and roll back.</p><p>The mechanics rhyme across all three. There's a set of data files (Parquet, usually) holding the actual rows. Above that sits a metadata layer that tracks which files belong to the table right now. In Iceberg's vocabulary that's manifest files listing data files plus their column-level stats, manifest lists grouping those manifests, and metadata files describing the table's current state. Every write produces a new snapshot - an immutable pointer to the exact set of data files that constituted the table at that moment. A commit is, at bottom, an atomic swap of "the current snapshot is now this one." Delta does the same thing with a transaction log (the <code>_delta_log</code> directory of ordered JSON commits); Hudi does it with timeline files and a slightly different file-grouping model tuned for upserts. Different nouns, same idea.</p><p>Once the snapshot model clicked for me, the headline features stopped looking like separate features and started looking like consequences of that one idea. ACID on object storage: a reader always sees a consistent snapshot because the snapshot is immutable and the commit is a single atomic pointer flip, so a reader either sees the old set of files or the new one, never a half-written mess. Time travel: old snapshots aren't deleted immediately, so "show me the table as of last Tuesday" is just reading an older pointer. Schema evolution: the metadata tracks column identity over time, so adding, renaming, or dropping a column is a metadata operation, not a rewrite of every file. None of this is magic. It's careful bookkeeping in a metadata layer that turns dumb object storage into something transactional.</p><h2>How Iceberg pulled ahead</h2><p>For a few years this was a genuine three-way contest with real religious wars attached. Delta Lake had Databricks behind it and the gravity that comes with being the default in the most popular Spark platform on earth. Hudi came out of Uber, built streaming-first, the one that took incremental upserts and change capture seriously when the others were batch-shaped. Iceberg came out of Netflix, designed by people who'd been burned by Hive's directory-listing model and wanted something that scaled to enormous tables without melting the metastore.</p><p>Iceberg pulled ahead on two things. One was technical: its design was the cleanest for very large tables and the most engine-neutral from the start - it never assumed Spark the way Delta effectively did. The other was political, and it mattered more. Iceberg's open governance under the Apache Foundation, plus a REST catalog spec that any engine could implement, made it the safe Switzerland choice for vendors who didn't want to hand Databricks the keys to their data layer. If you're Snowflake or AWS or Trino, adopting Delta means adopting Databricks' format; adopting Iceberg means adopting a neutral standard. That's not a hard call.</p><p>The decisive moment was Databricks buying Tabular in June 2024 for north of a billion dollars. Tabular was the company founded by Iceberg's original Netflix creators. Read that again: the company most identified with Delta Lake paid over a billion dollars to acquire the people who built the competing format. That is not what winning a format war looks like. That's the Delta camp acknowledging Iceberg's gravity and buying a seat at its table. Databricks' own framing was convergence - bring the Delta and Iceberg creators under one roof and work toward a single interoperable standard - and they were explicit that it'd take years, not months.</p><h2>Convergence, not conquest</h2><p>The thing that actually killed the war is that the formats stopped being mutually exclusive. Delta Lake UniForm is the clearest example. UniForm writes your data as Delta but also generates Iceberg (and Hudi) metadata pointing at the same underlying Parquet files, so an Iceberg reader can query a Delta table with no copy and no conversion. One set of data files, multiple metadata layers describing it, read it as whatever your engine speaks. The Parquet was never the disputed territory - it's the same columnar bytes regardless. The fight was always about the metadata layer on top, and once that layer became something you could project into multiple formats over shared files, "which format" stopped being a one-way door.</p><p>Iceberg's own evolution pushes the same direction. The format's newer spec revisions added the row-level delete and update machinery that used to be the reason you'd pick Hudi for mutation-heavy workloads, narrowing the gap that justified a separate streaming-first format. So you have Delta growing the ability to be read as Iceberg, and Iceberg growing the capabilities that were Hudi's whole pitch. The three formats are bleeding into each other on purpose.</p><p>The industry settled on a phrase for it: write once, read anywhere. Whether you can fully deliver that today depends on which features and which engine versions you're on - the interop is real but the edges are still rough, and "UniForm makes Delta readable as Iceberg" carries asterisks around exactly which Iceberg capabilities survive the projection. So treat write-once-read-anywhere as the clearly-stated direction of travel that's already mostly true for common cases, not as a finished guarantee for every feature.</p><h2>Where Hudi sits now</h2><p>Hudi isn't dead, whatever the migration stories suggest. It remains genuinely strong at what it was built for: high-frequency upserts, incremental pulls, streaming ingestion where records mutate constantly and you need merge-on-read and record-level indexes to keep write amplification sane. If your workload is a firehose of changing records rather than mostly-append analytics, Hudi's design still has real advantages, and there are large production deployments running on it happily.</p><p>But its center of gravity narrowed. It went from one of three contenders for the default open table format to a specialist tool you reach for when its particular streaming-upsert strengths matter more than ecosystem breadth. The honest read in 2026: greenfield projects mostly default to Iceberg, Hudi stays where the streaming-mutation profile justifies the smaller ecosystem, and a chunk of the teams that picked it in 2021-2022 are doing what that first team did - migrating, not because Hudi failed them but because being the odd format out has a tax that compounds.</p><h2>The catalog is the new battleground</h2><p>Here's the part that matters more than which format you write. Once everyone can read everyone's files, the format is commoditized, and the thing that isn't commoditized is the catalog - the service that knows which tables exist, where their current metadata lives, who's allowed to touch them, and how to atomically commit a new snapshot. The catalog is the control plane. It owns governance, access control, and the commit protocol. Whoever owns the catalog owns the actual lock-in.</p><p>Iceberg's REST catalog spec turned the catalog into a standard HTTP API: implement the spec and any engine that speaks it can talk to your catalog. That spec is now the contested ground, and the contestants are exactly who you'd expect. Snowflake built Polaris, an open-source Iceberg catalog, and donated it to the Apache Foundation in 2024 (partnering with Dremio, AWS, Google and Microsoft on the donation); Snowflake's managed version of it ships as Snowflake Open Catalog, GA since October 2024. Databricks open-sourced Unity Catalog and took the multi-format route - Unity aims to catalog Delta, Iceberg, Hudi, and unstructured files alike, rather than betting on Iceberg alone. AWS shipped S3 Tables, object storage with Iceberg and a catalog baked directly into the bucket. Every cloud and every warehouse now has a horse in the catalog race, and they all speak the Iceberg REST dialect to some degree.</p><p>That's the shape of the next few years. The format question is settled enough that vendors give the format away for free, because the format is no longer where the money or the lock-in lives. The catalog is. Format war over, catalog war beginning - and the catalog war is the one to actually pay attention to, because the catalog is where governance, security, and switching costs concentrate.</p><h2>How I'd actually choose in 2026</h2><p>The first thing I tell anyone agonizing over this is that the decision carries far less weight than it did three years ago. With that said, here's roughly how I reason through it:</p><ul><li><p>On greenfield with no overwhelming reason otherwise, I pick Iceberg. It's the de facto open standard, it has the broadest engine and vendor support, and choosing it means you're the normal case every tool plans for instead of the special case. The ecosystem tax runs in your favor.</p></li><li><p>If a team is already deep in the Databricks world, Delta is fine and UniForm is the bridge. There's no need to rip out Delta to interoperate - write it with UniForm enabled and Iceberg readers can query it. The convergence work exists precisely so this isn't a forced migration.</p></li><li><p>Heavy streaming upserts and mutation-dominated ingestion are where Hudi still earns its place. If merge-on-read and record-level indexing on a constant stream of changing records is your actual workload, its design advantages are real. Just go in clear-eyed that you're choosing the narrower ecosystem.</p></li><li><p>Whatever the format, the catalog is where I'd spend the real attention. Which one you adopt - Polaris/Open Catalog, Unity, a cloud-native one like S3 Tables, or a self-hosted REST catalog - determines your governance model, your access-control story, and how locked-in you are. The format increasingly takes care of itself through interop.</p></li><li><p>And I'd demand REST-catalog compatibility from anything I adopt. The Iceberg REST spec is the interop seam. A tool that speaks it slots into a multi-engine world; one that insists on its own proprietary catalog protocol is the thing that traps you. That's the lock-in axis that's still live, so guard it.</p></li></ul><h2>The ones I keep seeing</h2><ul><li><p>Re-litigating the format war like it's still 2022 is the one I run into most. Teams burn weeks in format-selection meetings as if the choice is load-bearing, when for most analytics workloads it stopped being load-bearing once interop landed. Default to Iceberg and move on to problems that still matter.</p></li><li><p>Then there's confusing the format with the catalog. "We use Iceberg" tells me your file metadata layout and nothing about who governs your tables or how locked in you are. Those questions live in the catalog, and conflating the two means skipping the decision that actually has consequences.</p></li><li><p>Assume "write once, read anywhere" is already a clean guarantee everywhere and you'll get bitten on the edges. The interop is real and improving fast, but UniForm projections and cross-format reads carry feature and version asterisks. Verify the specific capabilities you depend on survive the round trip before you architect as if they do.</p></li><li><p>Picking Hudi for batch analytics because it was once a contender is a recurring waste. Its strengths are streaming and upserts; on mostly-append analytics you're choosing the narrowest ecosystem for none of the benefit. Match the format to the actual write pattern.</p></li><li><p>Letting a vendor's proprietary catalog become your lock-in by default is the trap that's quietly replaced the old one. The format being open doesn't help if the catalog protocol is closed, and the catalog is where switching cost concentrates now.</p></li><li><p>Last, treating an open-source catalog and its managed cloud version as identical. Polaris the Apache project and Snowflake Open Catalog the managed service are related but not the same surface; Unity Catalog open-source and the Databricks-hosted Unity are likewise distinct. Know which one you're actually running and what its governance story is.</p></li></ul><p>The whole arc here is a commodity forming in real time. Five years ago the table format was a strategic, near-irreversible bet, and three vendors fought over it like it was the whole game. Then everyone realized the Parquet underneath was never in dispute, the metadata layer could be projected into multiple formats over the same files, and the engines could just learn to read all of them. So the format stopped being the moat. What's left as a moat is the catalog - the control plane that decides what exists, who can touch it, and how a write becomes official. If you take one thing from all of this, let it be where you point your scrutiny: not at which format to write, which is nearly settled, but at which catalog to trust, which is wide open and decides how free you'll be to change your mind later.</p>]]></content:encoded></item><item><title><![CDATA[kube-apiserver watch cache: the layer that saves etcd]]></title><description><![CDATA[LIST-WATCH, resourceVersion, bookmarks, relist storms, consistent reads]]></description><link>https://podostack.com/p/k8s-api-server-watch-cache</link><guid isPermaLink="false">https://podostack.com/p/k8s-api-server-watch-cache</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Wed, 01 Jul 2026 14:02:56 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!D6x4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!D6x4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!D6x4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!D6x4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!D6x4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!D6x4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!D6x4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!D6x4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!D6x4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!D6x4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8b77b06f-bdfa-4025-8fcd-3f18195d641c_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The page that night was apiserver OOMKilled, then OOMKilled again forty seconds after the kubelet brought it back. Three replicas, all dying in a rolling wave. The cluster was big - thousands of namespaces, north of two hundred thousand secrets - and somebody had shipped a controller that, on every reconcile, did a plain <code>client.List</code> of all secrets across all namespaces with no pagination and no informer. One pod restart kicked off a reconcile, the reconcile pulled the entire secret set into apiserver memory at once, the apiserver fell over, every other controller noticed the connection drop and relisted, and now you've got a thundering herd hammering the one component everything depends on. We spent the first twenty minutes assuming etcd was the problem. It wasn't. etcd was bored. The apiserver was the thing eating itself, and the watch cache is the layer that's supposed to stop exactly this.</p><h2>What a LIST actually costs</h2><p>People talk about "the Kubernetes API" like it's a database you query. It's closer to a cache in front of a database, and the database is etcd. Every object - every Pod, Secret, ConfigMap, CRD instance - is a key-value pair in etcd, serialized as protobuf. When you ask the apiserver for something, the interesting question is whether it answers from etcd or from its own memory.</p><p>A naive LIST straight against etcd is brutal in a way that isn't obvious from the one-line client call. First, it's a quorum read: etcd is a Raft cluster, and a linearizable read has to confirm with a majority of members that the leader hasn't been superseded before it hands back data. That's a network round trip across the etcd peers before a single byte comes back. Then comes the part that actually hurt us - the apiserver pulls every matching key out of etcd, deserializes each one from protobuf into a Go object, holds the whole set in memory, re-serializes it into whatever the client asked for (JSON, usually, which is bigger and slower), and ships it. List a hundred thousand secrets and the apiserver is briefly holding a hundred thousand decoded objects in the heap. Do that from three controllers at once and you've got the heap graph we were staring at.</p><p>So most reads don't go to etcd at all. They go to the watch cache.</p><h2>The cacher</h2><p>Inside the apiserver, each resource type gets its own cache, an internal component called the Cacher. There's a Cacher for Pods, one for Secrets, one for every CRD. Each one opens a single long-lived WATCH against etcd for its slice of the keyspace and just... keeps up. etcd streams every create, update, and delete as it happens, and the Cacher applies those events to an in-memory copy of the current state. It's a tail of the etcd write log, materialized into a live picture of "what exists right now."</p><p>That's the trick. When ten thousand informers across the cluster all want to know about Pods, they don't each open ten thousand watches against etcd. They watch the apiserver, the apiserver watches etcd once, and the one upstream watch fans out to everyone downstream. etcd sees a handful of connections instead of tens of thousands. The cache absorbs the multiplexing.</p><p>A LIST served from the cache skips the quorum read and the per-object etcd fetch entirely. The objects are already decoded and sitting in memory. The cache filters them by namespace and label selector, serializes the result, done. That's why <code>kubectl get pods</code> feels instant on a cluster where a direct etcd LIST would take a real chunk of a second.</p><h2>What resourceVersion actually is</h2><p>Every object in Kubernetes carries a <code>resourceVersion</code>, and I've misread it myself more than once. It is not a per-object version counter. It's an opaque token derived from etcd's global revision - a monotonically increasing number that ticks on every write to the entire keyspace. Edit one ConfigMap and its resourceVersion jumps to whatever etcd's global revision happens to be at that instant, which might be a number thousands higher than its previous value because thousands of unrelated writes happened in between. Treat it as a logical clock for the whole store, not a version of the thing you're holding.</p><p>This matters because resourceVersion is how the watch cache reasons about freshness. When a client LISTs and gets back a result at resourceVersion 8000, then opens a WATCH "since 8000," the apiserver knows precisely where to resume - replay every event after revision 8000. No gaps, no duplicates. When you pass <code>resourceVersion=0</code> on a LIST, you're telling the apiserver "I don't care about being current, give me whatever the cache has, even if it's slightly stale." That's the cheap path informers use on startup, and it's the right default for them. When you omit resourceVersion entirely, you're asking for a strongly consistent read, and historically that meant a quorum read against etcd. Hold that thought.</p><p>The failure mode hiding in here is <code>410 Gone</code>. The cache only holds a sliding window of recent history. If a watch tries to resume from a resourceVersion that's already fallen off the back of that window - because the watcher was disconnected too long, or etcd compacted that revision away - the apiserver can't replay the missing events and returns <code>410 Gone</code>. The client's only correct response is to throw away its cached state and LIST again from scratch. Which is fine for one client. It is not fine when a network blip disconnects every watcher in the cluster at once and they all relist together.</p><h2>Bookmarks, and why your watch goes quiet</h2><p>Here's a subtle one. Suppose your controller watches Pods, but it filters to a label selector that matches almost nothing - one Pod in a busy cluster where thousands of other Pods churn every minute. Your watch sits silent for an hour because none of your one Pod's events fired. Meanwhile etcd's global revision has climbed by tens of thousands. Your client still thinks it's at the revision from an hour ago. When it eventually has to reconnect, that revision is long gone, you get a <code>410</code>, and you relist.</p><p>Watch bookmarks fix this. The apiserver periodically sends a near-empty event - no object payload, just a bumped resourceVersion - down every watch stream, including the quiet ones. It's the server saying "nothing you care about changed, but you're now caught up to revision 47000, write that down." The client advances its resourceVersion without doing any work, so when it reconnects it asks to resume from a revision the cache still has. Bookmarks turned <code>410</code> storms from a regular annoyance into a rare event. They've been on by default for years now, and if you're writing a controller against client-go you get them for free.</p><h2>The relist storm</h2><p>Back to the incident. The thing that turned one bad controller into a cluster-wide outage was the relist. When the apiserver died, every informer's watch connection dropped. client-go's reflector, on a broken watch, does the only safe thing: it relists. So at the moment all three apiservers were already starved for memory, every controller, every kubelet, every operator in the cluster simultaneously fired a LIST to repopulate its informer. The survivors got buried. This is the apiserver equivalent of a cache stampede, and it's why apiserver restarts on large clusters are genuinely scary - the recovery traffic can be worse than whatever knocked it down.</p><p><code>kubectl get pods -A</code> is a smaller version of the same thing. On a laptop against a dev cluster it's nothing. On a cluster with hundreds of thousands of Pods it tells the apiserver to assemble every Pod object, serialize the lot, and stream it to your terminal - a multi-hundred-megabyte response materialized in the apiserver heap for one human running one command. Run it during an incident, when the apiserver is already under pressure, and you can be the one who tips it over. I've watched it happen. The <code>-A</code> is doing a lot of quiet work.</p><h2>The 1.5 MB wall and why pagination exists</h2><p>etcd refuses any request above roughly 1.5 MB by default - that's the <code>--max-request-bytes</code> limit, and it's there for a reason. A single value can't be allowed to grow unbounded or it threatens the Raft log. The apiserver respects it. But a LIST result is obviously bigger than 1.5 MB on any real cluster, so the apiserver doesn't fetch the whole thing in one etcd call - it pages through etcd's keyspace in chunks under the hood.</p><p>Pagination is also exposed to you, and you should use it. A LIST with <code>limit=500</code> returns 500 items plus a <code>continue</code> token, an opaque cursor pointing at where to resume. You pass it back on the next request and walk the collection in bounded pieces. The win isn't on the wire so much as in memory: the apiserver allocates a buffer for 500 objects, not for two hundred thousand. The unpaginated controller in our incident was the entire problem in one sentence - it asked for everything at once and the apiserver tried to give it everything at once.</p><p>There's a wrinkle worth knowing. Paginated LISTs historically could not be served from the watch cache - the cache held current state, not the snapshot-at-a-revision you need to page consistently, so a paginated LIST would fall through to etcd. For years that put you in a bind: paginate and hit etcd, or list-from-cache and blow up memory. Recent Kubernetes closed that gap, which is the next part.</p><h2>Serving consistent reads from cache</h2><p>For a long time the rule was simple and annoying: if you wanted a strongly consistent read - the latest data, guaranteed - it had to come from etcd via a quorum read, because the cache might be a few milliseconds behind. Since more than 80% of LISTs the apiserver handles are these consistent reads, that's a lot of load landing on etcd for data the cache very nearly already had.</p><p>The Consistent Reads from Cache work changed the deal. It leans on etcd's progress notifications: etcd periodically tells each watcher "you're current as of revision N." The apiserver records the latest revision it's seen from etcd, and when a consistent LIST comes in, it waits until the cache has caught up to at least that revision - a freshness check with a short timeout - then serves from memory. Same consistency guarantee, no per-object etcd fetch. The published numbers are good: in 99.9% of cases the cache became fresh enough within about 110 milliseconds. This landed as beta and on by default back in v1.31 (August 2024), and as of mid-2026 it's still riding as beta rather than formally GA, so check your version's feature gates if you're depending on the behavior.</p><p>Then there's the snapshot work. The Snapshottable API server cache (beta and default-on as of v1.34, September 2025) gives the cache lightweight, point-in-time snapshots - lazy copies built on a btree, sharing object pointers rather than duplicating data - so historical-revision and paginated LISTs can finally be served from cache instead of falling through to etcd. The bind I mentioned above mostly goes away: you can paginate and stay in the cache.</p><p>The streaming side moves too. WatchList, also called streaming lists, replaces the relist-then-buffer pattern with a watch-based initial list that streams objects one at a time instead of assembling the whole collection in memory first. The reported effect is dramatic - apiserver memory stabilizing around 2 GB in a test where the old path climbed to roughly 20 GB. The rollout has been bumpy: enabled by default in v1.32, then reverted to off-by-default in v1.33, and the client side still needs you to opt in via <code>WatchListClient</code>. So it's real and it's coming, but it is not yet the thing quietly protecting your cluster unless you turned it on.</p><h2>Why informers exist at all</h2><p>All of this is the backdrop for the single most important client-side fact: do not poll the apiserver, watch it. An informer (client-go's cache, the thing controller-runtime wraps) does one LIST to seed itself, then holds a long-lived WATCH and keeps a local copy of every object it cares about. After that initial list, your controller reads from its own in-process memory at zero cost to the apiserver. A reconcile loop that hits a shared informer cache makes no API calls; a reconcile loop that calls <code>client.List</code> every time makes one expensive call every time. Our incident controller did the second thing, against all secrets, with no pagination. That's three separate mistakes stacked into one reconcile, and any one of them removed would have kept the cluster up.</p><h2>The decision framework</h2><p>When you're touching the read path on a cluster of any real size, walk through this honestly.</p><ul><li><p>An informer comes first, always. If you're reconciling, you want a shared informer or controller-runtime client backed by a cache, not raw <code>List</code>/<code>Get</code> per loop. The informer LISTs once and watches forever, which is the difference between zero apiserver calls per reconcile and one heavy one.</p></li><li><p>Then I scope the watch. A label selector or a single-namespace watch means the cache filters down to what you need and your event stream stays small. Watch all objects of a type cluster-wide when you care about a handful, and you end up watching everything and bookmarking constantly.</p></li><li><p>Anything outside an informer - a CLI tool, a migration script, an audit job - gets pagination: set a <code>limit</code>, walk the <code>continue</code> token, never ask for the unbounded set. The buffer the apiserver allocates is proportional to your limit, not the collection size.</p></li><li><p>Where stale is fine - informer warm-up, dashboards, anything that doesn't need this-instant accuracy - <code>resourceVersion=0</code> takes the cheap cache read. Consistent reads are worth saving for when correctness actually depends on freshness.</p></li><li><p>And before leaning on any of the cache behavior, check the cluster's version. Consistent-reads-from-cache, snapshottable cache, and WatchList have all moved through alpha/beta/reverted states on different timelines, so the protection you assume is on might not be.</p></li></ul><h2>Common mistakes</h2><ul><li><p>Unpaginated cluster-wide LISTs in a hot path. The classic. One <code>client.List</code> of everything, called on every reconcile, against a resource with a lot of objects. It works in staging with fifty objects and detonates in production with two hundred thousand.</p></li><li><p>Polling instead of watching. A controller that GETs or LISTs on a timer when it should hold a watch. Every poll is load the informer would have eliminated; you're paying for a database query to learn nothing changed.</p></li><li><p>Treating <code>resourceVersion</code> as a per-object counter. It's a global etcd revision. Comparing two objects' resourceVersions to decide which is "newer" within the same object is fine; reasoning about the absolute number as if it counts edits to that one object is wrong and will mislead you.</p></li><li><p>Ignoring <code>410 Gone</code> instead of relisting. The only correct response to a <code>410</code> is to drop cached state and list again. Retrying the same watch from the same dead revision just loops. client-go handles this for you - hand-rolled watch code often doesn't.</p></li><li><p><code>kubectl get pods -A</code> during an incident. On a big cluster that's a multi-hundred-megabyte response assembled in an already-stressed apiserver heap. Scope it to a namespace, add a selector, or paginate, especially when things are on fire.</p></li><li><p>Per-replica watches that don't share a cache. Multiple informers for the same resource inside one process, each opening its own watch, when a shared informer factory would open one. The apiserver fans out cheaply, but you can still multiply your own footprint pointlessly.</p></li><li><p>Assuming a big cluster's apiserver restart is routine. It triggers a relist from everything at once. The recovery storm can be worse than the original fault, and that's the moment to have pagination and informers already in place, not to discover you don't.</p></li></ul><p>The whole architecture is one cache standing between a quorum-consistent log and ten thousand things that want to read it constantly. etcd holds the truth, the watch cache holds a live copy so etcd doesn't get mobbed, and informers hold a copy of the copy so the apiserver doesn't get mobbed either. Every outage I've seen on this path was someone reaching past one of those layers - hitting the apiserver like it was etcd, or hitting etcd like it was a local map. The layers exist because the thing underneath them can't survive being read directly at scale.</p>]]></content:encoded></item><item><title><![CDATA[Issue #024 - eBPF reads your TLS traffic without the private key]]></title><description><![CDATA[SSL_read uprobes, Go crypto/tls offsets, Pixie vs Beyla vs eCapture, kTLS, the threat model]]></description><link>https://podostack.com/p/issue-024-ebpf-reads-tls-without-keys</link><guid isPermaLink="false">https://podostack.com/p/issue-024-ebpf-reads-tls-without-keys</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Tue, 30 Jun 2026 14:01:42 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!YYyn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!YYyn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!YYyn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!YYyn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!YYyn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!YYyn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!YYyn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!YYyn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!YYyn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!YYyn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1427f32e-596a-45ff-a770-62e8bfb26827_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I was sitting in a Pixie session, scoped to one of our own namespaces, expecting the usual latency histograms and the boring shape of a healthy service. Then I clicked into a single request and the body was right there. The full JSON payload our payments service had POSTed to an internal API, headers and bearer token included, rendered as plaintext in a browser tab. The service talks TLS to everything. The mesh enforces mTLS. I had not handed Pixie a key, a cert, or a secret of any kind. It had been running for about four minutes.</p><p>The jolt wasn't "is this a bug." I knew roughly how it worked. The jolt was the gap between knowing it in the abstract and watching my own service's secrets scroll by in a tool I'd installed myself that afternoon. "Encrypted in transit" had quietly stopped meaning what I'd been letting it mean in my head, and it took seeing the payload to notice I'd been rounding it off.</p><p>This is about a technique that's load-bearing under half the observability stack you might be running, and is also, looked at from a slightly different angle, a clean way for someone with a foothold on your node to read everything your "encrypted" services say. Same mechanism. The only thing that changes is who's holding it.</p><h2>&#128300; The eBPF Trace: a uprobe on SSL_read sees plaintext before the key touches it</h2><p>For a long time I carried the same wrong picture of TLS observability that most people do. To read encrypted traffic you need a key - the session key, or the server's private key - or failing that you stand up a MITM proxy with a cert the client trusts. All of that is true if you're attacking the bytes on the wire. None of it is what these tools do.</p><p>The trick is to not touch the wire at all. Your application calls <code>SSL_write</code> with a pointer to a buffer. That buffer holds plaintext. OpenSSL takes it, encrypts it, and only then hands ciphertext to the kernel's <code>send()</code>. The encryption happens <em>inside</em> the library call, after your plaintext is already sitting in addressable memory. So if you can read that memory at the moment the function is entered, you get the plaintext, and you never have to know a single thing about the key the library is about to use. The same logic runs in reverse for reads: <code>SSL_read</code> returns decrypted plaintext into a caller-supplied buffer, so you wait for the function to return and read the buffer then.</p><p>eBPF gives you exactly that hook. A uprobe is a kernel-managed breakpoint on a userspace function. You attach one to <code>SSL_write</code> in <code>libssl.so</code>, and every time any process linked against that library calls it, your tiny eBPF program runs first, with the function's arguments in registers.</p><h3>The two-probe dance</h3><p>There's a wrinkle that's worth slowing down on, because it explains the shape of every tool that does this. For a write, the plaintext buffer is an <em>argument</em> - it exists when the function is entered. For a read, the plaintext doesn't exist yet on entry; the buffer is empty and the byte count isn't known until the library has decrypted and returned. So you can't capture a read with a single probe.</p><p>What the tools do is split it across an entry probe and a return probe. On <code>SSL_read</code> entry, the eBPF program stashes the buffer pointer in a BPF hash map, keyed by <code>bpf_get_current_pid_tgid()</code> - the 64-bit value that packs the process and thread ID. On <code>SSL_read</code> return, a second probe pulls that pointer back out of the map, reads the function's return value to learn how many bytes actually landed, and only then copies the buffer out. The pid_tgid key is what keeps two threads calling <code>SSL_read</code> at the same time from stealing each other's buffers. The copy itself is a <code>bpf_probe_read_user()</code> into a perf buffer or ring buffer, which is the channel the userspace agent drains asynchronously. None of this is exotic; it's the standard entry-map-return pattern you'd use for any function whose output you want.</p><p>People worry this is expensive, and the surprise is how cheap it is. The probe runs in-kernel, off the application's critical path - the app calls <code>SSL_write</code> and returns exactly as it always did; the eBPF program runs alongside. Numbers I've seen quoted put the constant per-event eBPF cost around 0.2 microseconds, with the write-side uprobe landing near 0.007% CPU and the read-side return probe closer to 0.3% on a busy process. I haven't benchmarked those figures myself, so treat them as the right order of magnitude rather than gospel, but they match the lived experience: I ran Pixie across a fleet for weeks and never saw it move a latency graph. The cost isn't CPU. The cost is the capability you handed out to get it, which is the whole second half of this issue.</p><p>A stripped-down version of the write side, in bpftrace pseudocode, reads about like this:</p><pre><code>uprobe:/usr/lib/x86_64-linux-gnu/libssl.so.3:SSL_write {
    // arg0 = SSL*, arg1 = buf, arg2 = num
    printf("pid %d wrote %d bytes:\n%r\n", pid, arg2, buf(arg1, arg2));
}</code></pre><p>That's the whole idea in four lines: hook the function, the buffer is <code>arg1</code>, the length is <code>arg2</code>, print it. A real tool adds the read-side return probe, the per-thread map, connection correlation, and protocol parsing on top. But the part that reads your plaintext is genuinely that small.</p><h3>Stitching it back to a connection</h3><p>The first time I looked at raw uprobe output it was a blob of JSON with no idea which socket it went out on, which isn't observability yet. The connection metadata - source, destination, which HTTP request this was - lives in the syscall layer, not in <code>SSL_write</code>. So the better tools also probe the <code>send</code>/<code>recv</code> syscalls and correlate. Pixie's design leans on a neat property here: when <code>SSL_write</code> is on the call stack, the underlying <code>send()</code> is usually right below it on the same stack, so the uprobe can reach down and grab the socket file descriptor that the plaintext is about to flow through. That's how the plaintext gets matched to a real five-tuple without the agent ever parsing a TLS record.</p><h3>Where the buffer stops being readable</h3><p>The technique has a hard floor, and it's worth knowing because it's the one thing that makes "just use eBPF" not universally true. If the application uses <strong>kernel TLS (kTLS)</strong>, the encryption moves out of the userspace library and into the kernel's record layer. The library hands plaintext to the kernel and the kernel encrypts it. A uprobe on <code>SSL_write</code> may still see the plaintext on the way in, but tooling that hooks lower, expecting userspace encryption, goes blind - it sees ciphertext or nothing useful. kTLS is still uncommon in typical app stacks, but it's the clean architectural answer to this whole class of introspection, and it's why the technique is a property of <em>how your app does TLS</em>, not a law of physics.</p><h3>Links</h3><ul><li><p><a href="https://blog.px.dev/ebpf-openssl-tracing/">Pixie: tracing SSL/TLS connections with eBPF</a></p></li><li><p><a href="https://eunomia.dev/tutorials/30-sslsniff/">eunomia: capturing SSL/TLS plaintext using uprobe</a></p></li><li><p><a href="https://docs.kernel.org/networking/tls-offload.html">Kernel TLS offload (kernel.org)</a></p></li><li><p><a href="https://oneuptime.com/blog/post/2026-01-07-ebpf-ssl-tls-inspection/view">oneuptime: inspecting SSL/TLS traffic with eBPF</a></p></li></ul><h2>&#127386; Showdown: Pixie vs Beyla vs eCapture, and what Go does to all three</h2><p>Once you know the mechanism, the tools sort themselves by two questions: what do they do with the plaintext once they have it, and how hard do they work to handle the cases where <code>libssl.so</code> isn't sitting there waiting to be probed. The second question is mostly a single word: Go.</p><h3>OpenSSL, the path nobody trips on</h3><p>For a process dynamically linked against OpenSSL, all three tools do roughly what the last section described. You point them at <code>libssl.so</code>, they find <code>SSL_read</code> and <code>SSL_write</code> as dynamic symbols, attach the entry and return probes, and the plaintext flows. GnuTLS and NSS work the same way with different function names (<code>gnutls_record_send</code>, <code>PR_Write</code>). This is the well-trodden path and it's been production-stable for years.</p><p>The thing that breaks it isn't the library, it's the <em>linking</em>. Statically linked OpenSSL has no <code>libssl.so</code> to attach to - the symbols are buried inside the application binary, and if it's been stripped, they may not be there to find by name at all. BoringSSL, which is what Envoy and a lot of Google-lineage software use, is almost always statically linked, so the same problem shows up there. The tool has to locate the function inside the host binary by symbol or by offset rather than by opening a shared object.</p><h3>Go breaks the model on purpose</h3><p>Go is the interesting case, and it's where the tools genuinely diverge. Go doesn't link OpenSSL. It ships its own TLS stack in the <code>crypto/tls</code> package and statically links it into every binary, because static linking is the whole Go distribution story. So there is no <code>libssl.so</code> anywhere. To read a Go service's TLS plaintext you have to probe <code>crypto/tls.(*Conn).Read</code> and <code>crypto/tls.(*Conn).Write</code> directly inside the application binary, located by symbol offset.</p><p>Then it gets worse, in two specific ways. First, since Go 1.17 the compiler passes function arguments in registers (the internal ABI), not on the stack the way the old convention did - so the eBPF program has to know which ABI the binary was built with to find the buffer pointer in the right place. Second, and this is the sharp one: a normal <code>uretprobe</code> is unreliable on Go. Go's runtime moves goroutine stacks around as they grow, and the kernel's return-probe trampoline doesn't survive that. The workaround the tools converge on is to disassemble the target function, find every <code>RET</code> instruction, and attach a plain <code>uprobe</code> at each of those offsets - manually reconstructing the return probe the kernel can't safely give you. It works, but it's per-binary, offset-sensitive, and a category of fragile that the OpenSSL path just doesn't have. The practical fallout: a Go service that recompiles with a new toolchain can shift those offsets, and a tool that cached them silently starts capturing garbage until it re-resolves. I lost an afternoon once to exactly that shape - a Go upgrade landed, the traces for that one service went empty, and nothing in the agent's logs said "your offsets are stale" because from its side nothing had failed. The probe was attached. It was just attached to the wrong place in a binary that had moved underneath it.</p><h3>Where each tool lands</h3><p>The one I had open in the intro was Pixie, which sits at the full-platform end. It traces both OpenSSL and Go's <code>crypto/tls</code>, feeds the captured plaintext through the same protocol parsers and perf buffers it uses for cleartext traffic, and hands back the L7 service map, latency, and request bodies in a UI. Now under CNCF, New Relic-backed. The plaintext capture is a feature you mostly don't think about; it just shows you the requests.</p><p>Grafana's Beyla I'd put a tier toward the metrics side. It got donated to OpenTelemetry in May 2025 and became the basis for OpenTelemetry eBPF Instrumentation (OBI). The framing is zero-code auto-instrumentation: it taps <code>libssl</code> to pull HTTP-level information out of SSL traffic, and you get OTel metrics and traces out the other end without touching the app. Beyla leans hard on the Go case, because Go is where "I can't add an SDK easily" hurts most. One wrinkle I keep having to remind people of: to propagate trace context across encrypted HTTP it wants to be present on both ends and reaches for Linux Traffic Control at the packet level, which is a different mechanism bolted on top of the uprobe capture.</p><p>Then there's eCapture, which drops the observability pretense and just hands you the bytes. A CLI, around 15k GitHub stars, supporting OpenSSL, GnuTLS, NSS, BoringSSL, and Go. Its three modes tell you exactly what it's for: text mode prints plaintext to your terminal, pcap mode writes a Wireshark-openable capture of the <em>decrypted</em> traffic, and keylog mode dumps the TLS master secrets so you can decrypt a capture after the fact. No agent, no UI, no metrics pipeline. Point it at a process, read its TLS. Root or <code>CAP_BPF</code>, kernel 4.18+ on x86_64, and you're done.</p><p>One name I keep having to pull out of this list is Cilium. Its TLS visibility looks like it belongs here and it doesn't - it terminates the connection at Envoy with a cert your workload trusts, inspects, and re-originates. That's classic key-based MITM, not uprobe-on-plaintext. Different threat model, different requirements, different failure modes. If someone tells you "Cilium reads TLS the same way Pixie does," they're wrong about both.</p><h3>Links</h3><ul><li><p><a href="https://blog.px.dev/ebpf-tls-tracing-past-present-future/">Pixie: eBPF TLS tracing, past, present and future</a></p></li><li><p><a href="https://speedscale.com/blog/ebpf-go-design-notes-1/">Speedscale: under the hood with Go TLS and eBPF</a></p></li><li><p><a href="https://github.com/gojue/ecapture">eCapture on GitHub</a></p></li><li><p><a href="https://grafana.com/blog/2025/05/07/opentelemetry-ebpf-instrumentation-beyla-donation/">Grafana: why we donated Beyla to OpenTelemetry (OBI)</a></p></li><li><p><a href="https://docs.cilium.io/en/stable/security/tls-visibility/">Cilium: inspecting TLS encrypted connections</a></p></li></ul><h2>&#128293; Hot Take: encrypted in transit is a promise about the wire, not the node</h2><p>Here's the part I sat with longest after the Pixie session. The capability that gives my observability stack its L7 view is, byte for byte, the capability an attacker uses to read my "encrypted" traffic. There is no second technique. eCapture in text mode and Pixie's request inspector are the same uprobe on the same <code>SSL_write</code>, pointed at the same buffer. One of them ships me a flame graph and one of them ships me your bearer tokens, and the only difference is intent and who ran it.</p><p>That collapses a mental model a lot of us carry without examining it. "TLS everywhere," "mTLS in the mesh," "encrypted in transit" - we say these like they're a guarantee of confidentiality. They're not. They're a guarantee about the <em>wire</em>. They protect the bytes once they leave the process and right up until they enter the next one. Inside the process, before <code>SSL_write</code> encrypts and after <code>SSL_read</code> decrypts, the data is plaintext in memory, and anyone who can run an eBPF program in that pid namespace can read it without breaking a single cryptographic primitive. The crypto is fine. The crypto was never the question. The question is who has code execution next to your process.</p><p>So the honest version of the security boundary is narrower than the marketing. To pull this off, an attacker needs <code>CAP_BPF</code> (or <code>CAP_SYS_ADMIN</code> on older kernels) and reach into the target's pid namespace - which in practice means a privileged container, a node compromise, a <code>hostPID: true</code> pod, a debug sidecar nobody locked down, or a DaemonSet with more capability than it needed. That's a real bar. It is not the bar most people picture when they picture "the traffic is encrypted." The traffic being encrypted does nothing against an adversary who already stands where your observability agent stands. And here's the uncomfortable symmetry: if you've deployed Pixie or Beyla or any eBPF L7 tool as a DaemonSet, you have already granted exactly that capability, to that agent, on every node. You decided this was fine. You were probably right. But you decided it, and the decision was "a privileged process may read every service's plaintext," which is a larger decision than "let's get better traces."</p><p>The reframe I landed on is to stop treating transit encryption and node trust as one property. They're two. TLS handles the first and handles it well. The second - who is allowed to execute privileged code beside your workloads - is a separate control plane: capability hygiene, <code>CAP_BPF</code> as a guarded grant, locking down <code>hostPID</code> and privileged pods, watching what your DaemonSets actually run. In the postmortems I've read, the teams that got burned weren't running broken TLS at all. They'd let "encrypted in transit" quietly stand in for "confidential on the box," and then handed broad eBPF capability to three different agents because each one's install doc made it sound routine. The wire was never the soft part. The node was.</p><h2>Until next week</h2><p>What stuck with me from the Pixie session wasn't the technical surprise. It was realizing how much weight I'd been putting on a phrase - "it's encrypted" - that was only ever true about one segment of the path. The bytes were safe in flight and fully exposed at rest in memory, and I'd been mentally filing both under the same word.</p><p>Next Tuesday, the thing this keeps circling back to: who holds the secret. Specifically, the four wrong ways to put secrets in Git, and how SOPS, Sealed Secrets, External Secrets, and Vault each draw the line between "in the repo" and "in the cluster" differently. If this issue made you slightly paranoid about what's readable on your nodes, that one is the natural next paranoid afternoon. See you then.</p><p>- Ilia</p>]]></content:encoded></item><item><title><![CDATA[Postgres partitioning: the pruning that doesn't happen]]></title><description><![CDATA[Declarative partitioning, partition pruning, partition-wise joins, Citus, sharding]]></description><link>https://podostack.com/p/postgres-partitioning-native-vs-citus</link><guid isPermaLink="false">https://podostack.com/p/postgres-partitioning-native-vs-citus</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Fri, 26 Jun 2026 14:01:52 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ThWF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ThWF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!ThWF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!ThWF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!ThWF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ThWF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ThWF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!ThWF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!ThWF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!ThWF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a0e2f45-3b5a-426c-8e49-52b0e6c92789_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The table had been "partitioned for performance" six months before I got there, and the queries were slower than before anyone touched it. Monthly range partitions, two years of them, twenty-four child tables under one events table. Looked textbook. Then I read the slow query log and the thing every report ran was a filter on <code>user_id</code> and a date range expressed as <code>created_at &gt;= now() - interval '7 days'</code>. The partition key was <code>event_date</code>. Not <code>created_at</code>, a different column the ETL had stopped populating consistently. So every query scanned all twenty-four partitions, every time, and paid the planning overhead of twenty-four tables on top. Partitioning hadn't sped anything up. It had added a tax and pruned nothing, because the planner never had the key it needed to prune.</p><p>That's the failure I want to talk you out of. Partitioning is a manageability tool that can speed some queries up as a side effect. It is not a performance feature you turn on. When it makes things faster it's because of pruning, and pruning only happens under conditions that are easy to miss.</p><h2>What declarative partitioning actually gives you</h2><p>Since Postgres 10 you write <code>PARTITION BY</code> and the database maintains the routing. There are three strategies. RANGE splits on ordered values, dates and sequential IDs being the common cases. LIST splits on discrete values, like a region or a tenant. HASH spreads rows across a fixed number of partitions by a hash of the key, for when you want even distribution and have no natural range. You declare the parent, attach child partitions, and inserts land in the right child automatically.</p><p>What you have not done is distribute anything. Every partition lives in the same Postgres instance, on the same disk, served by the same process. Native partitioning is organization, not scale-out. It's one table the planner is allowed to treat as several. That distinction is the whole back half of this piece, so hold onto it.</p><p>The reason to do it at all comes down to three things that get genuinely easier. Dropping old data becomes a <code>DETACH</code> and a <code>DROP</code> of one partition instead of a <code>DELETE</code> that bloats the heap and chews autovacuum. Bulk loads can target one partition. And the planner can sometimes skip partitions entirely. That last one is pruning, and it's the only one that touches query speed.</p><h2>Pruning, and when the planner can't do it</h2><p>Partition pruning is the planner deciding it doesn't need to look at a partition because the query's conditions can't match anything in it. Filter on the partition key with a value that falls in March, and Postgres reads the March partition and ignores the other twenty-three. That's the speedup people are chasing. It's real, and it's worth a lot when it fires.</p><p>The catch is in the word "key". Pruning needs the partition key in the query's conditions. Omit it, and there's nothing to prune on, so the planner reads every partition. A query that filters on <code>user_id</code> against a table partitioned by date will scan all of them and then filter inside each, which is strictly more work than the same query against an unpartitioned table with the right index. This is exactly what bit the table I opened with. Partitioning a table on a column your queries don't filter on is worse than not partitioning it.</p><p>There are two moments pruning can happen, and the difference matters when you read plans. Plan-time pruning is when the partition key is a constant the planner sees while planning, like a literal date. It drops the partitions right there and the plan only mentions the survivors. Runtime pruning is when the key's value isn't known until the query runs, which covers parameters in a prepared statement and the inner side of a nested loop join. Postgres can still prune in those cases, but it happens during execution, so <code>EXPLAIN</code> without <code>ANALYZE</code> will show all the partitions in the plan even though most get skipped at run time. I've watched people panic at a plan listing forty partition scans that, when actually run, touched two. Run it with <code>ANALYZE</code> and look at <code>(never executed)</code> on the partitions that got pruned.</p><p>One more way pruning quietly fails: wrap the partition key in a function or a type cast and the planner often loses the ability to prune. A query that filters on <code>date_trunc('day', event_date)</code> instead of plain <code>event_date</code> can scan every partition, because the planner reasons about the bare column, not the expression around it. Same with an implicit cast when the literal type doesn't match the column type. The query looks like it filters on the key, the plan disagrees, and the only way to catch it is to read the plan rather than trust the SQL.</p><p>Older Postgres had <code>constraint_exclusion</code>, a cruder mechanism from the inheritance-based partitioning era. It compares CHECK constraints against the query at plan time only, it's controlled by a GUC that defaults to <code>partition</code>, and it can't do runtime pruning at all. If you're on modern declarative partitioning you mostly get the better pruning automatically, but mixed setups and legacy inheritance trees can still fall back to constraint exclusion and its plan-time-only limits, which is one more reason a plan shows partitions you expected gone.</p><h2>The cost of too many partitions</h2><p>The instinct after learning pruning is to partition finely: daily partitions, a partition per tenant, thousands of them. That instinct has a sharp edge.</p><p>Planning time grows with partition count. The planner has to consider each partition, and even with pruning it does work proportional to how many it started with before pruning. A query against a table with a few dozen partitions is fine. The same query against several thousand can spend more time planning than executing, especially for short OLTP queries where the execution was going to be a millisecond anyway. There's a real ceiling here, and it's lower than people expect.</p><p>Locks are the other tax. Operations that need to touch the parent can acquire locks across all partitions, and a query that can't prune locks every one of them. Under concurrency that's a lot of lock entries, and the lock table is finite. I've seen a cron job that scanned an over-partitioned table blow past <code>max_locks_per_transaction</code> and take down writes that had nothing to do with it. And every partition is its own table with its own statistics, its own autovacuum bookkeeping, its own indexes to maintain. A thousand partitions is a thousand of each. The manageability win you partitioned for starts eating itself.</p><h2>Partition-wise joins, and the flag that's off</h2><p>When you join two tables partitioned the same way on the same key, Postgres can join them partition by partition instead of joining the whole things and sorting it out after. March-against-March, April-against-April, never crossing. For large partitioned joins that's a big saving, and the same idea applies to aggregates that group by the partition key.</p><p>The catch that surprises people: it's off by default. <code>enable_partitionwise_join</code> and <code>enable_partitionwise_aggregate</code> both default to off, because the planning cost is higher and it only pays off when the partitions line up. So you can build two perfectly matched partitioned tables, join them, and get the naive plan because nobody flipped the GUC. The feature exists, you opted into the partitioning, and you still don't get the optimization until you ask for it explicitly.</p><h2>Attach, detach, and rolling windows</h2><p>The pattern that justifies partitioning for a lot of teams is the time window. You keep ninety days of events, you partition by day or week, and aging out old data is a <code>DETACH PARTITION</code> followed by a <code>DROP TABLE</code> on the detached child. No giant <code>DELETE</code>, no bloat, no vacuum storm. The space comes back instantly because you dropped a whole table.</p><p>Going the other way, <code>ATTACH PARTITION</code> folds a populated table into the parent. Postgres validates that its rows fall within the partition bounds, which is a scan unless you've pre-created a matching CHECK constraint that lets it skip the validation. On a big table that's the difference between an attach that takes a lock for a second and one that scans for minutes. Worth knowing before you run it in the middle of the day. Detach got a <code>CONCURRENTLY</code> option in Postgres 14 specifically so you can pull a partition out without a long lock on the parent, which matters when the parent is serving live traffic.</p><h2>Indexes on partitioned tables</h2><p>Create an index on a partitioned parent and Postgres creates a matching index on every partition and on every partition you add later. That's usually what you want, and it mostly behaves like a normal index per partition. The corner is uniqueness. A unique constraint on a partitioned table has to include the partition key, because Postgres can only enforce uniqueness within a partition, not across them. So you can't declare <code>UNIQUE (email)</code> on a table partitioned by region the way you would on a plain table. The constraint becomes <code>UNIQUE (email, region)</code>, which is not the same guarantee. People discover this when their dedup logic lets the same email through in two regions, and the fix is a different schema, not a different index.</p><h2>Where native ends and Citus begins</h2><p>The word "partitioning" muddies this one, because Citus calls its thing partitioning too and it's a different animal. Native partitioning organizes one table on one machine. Citus distributes one logical table across many machines, with a coordinator that routes and a set of workers that each hold shards. Whenever a team asks me "which is better," I push back: the real fork is whether your problem is organization or capacity, and the two tools sit on opposite sides of it.</p><p>Native is enough when your data fits one machine and your pain is manageability or a specific pruning win. A few hundred gigabytes, time-series you age out, a hot date range your queries actually filter on, all single-node problems. Adding Citus there buys you a coordinator, a multi-node cluster, network hops on every query, and a distribution key you have to design your schema around, in exchange for solving a problem you didn't have. I've watched a team reach for Citus at 200 GB and spend a quarter on operational overhead that a few range partitions and an index would have settled in an afternoon.</p><p>You actually need Citus when one machine genuinely can't hold the data or serve the write throughput, and when your access pattern has a distribution key that most queries carry, so the coordinator can route to a single worker instead of fanning out to all of them. Multi-tenant SaaS where every query filters by <code>tenant_id</code> is the textbook fit, because the tenant key shards cleanly and most queries hit one shard. Get the distribution key wrong and you're back to fan-out across every worker, which is the distributed version of the no-pruning problem from the start of this piece, only now it's spread over the network. The failure mode rhymes. A query without the key scans everything, whether everything is twenty-four local partitions or forty remote shards.</p><h2>The ones I keep running into</h2><p>These show up again and again, roughly in the order they tend to bite:</p><ul><li><p>Partitioning on a column your queries don't filter on. No pruning, full fan-out, plus the per-partition planning tax. It's the single most common way partitioning makes things slower, and it's exactly what I walked into in the opener.</p></li><li><p>Then the <code>EXPLAIN</code>-without-<code>ANALYZE</code> panic at every partition in the plan. Runtime pruning doesn't show up there. Run it for real and check for <code>(never executed)</code>.</p></li><li><p>Partitioning too finely is the next trap: thousands of daily partitions turn planning time and lock counts into the bottleneck, and a cron job over the parent can exhaust the lock table and take writes down with it.</p></li><li><p>Expect partition-wise joins to just work and you'll be disappointed - they're behind <code>enable_partitionwise_join</code>, off by default, so matched partitioned tables still get the naive plan until you flip it.</p></li><li><p>A unique constraint doesn't behave like it does on a plain table either. It has to include the partition key, so <code>UNIQUE (email)</code> quietly becomes per-partition uniqueness and your dedup leaks across partitions.</p></li><li><p>Reaching for Citus before you've hit a single-machine wall is the expensive one. If the data fits one box and the queries carry the right key for native pruning, Citus is operational weight you're carrying for nothing.</p></li><li><p>And last, a Citus distribution key that most queries don't include. Every query fans out to every worker, the same no-pruning failure as native partitioning, now distributed across the network and far harder to debug.</p></li></ul><p>Partitioning earns its keep on manageability: dropping old data cleanly, loading in bulk, keeping a rolling window without a vacuum storm. The query speedup is a conditional bonus that shows up only when the planner can prune, which means only when your queries carry the partition key. Decide what you're actually buying before you cut the table into pieces. Most of the time the honest answer is "easier data lifecycle", and if you went in expecting "faster queries" you'll end up like that events table, paying a tax and pruning nothing.</p>]]></content:encoded></item><item><title><![CDATA[Redis persistence: why RDB and AOF both lie about durability]]></title><description><![CDATA[RDB snapshots, AOF fsync, fork latency, hybrid preamble, replication]]></description><link>https://podostack.com/p/redis-persistence-rdb-vs-aof</link><guid isPermaLink="false">https://podostack.com/p/redis-persistence-rdb-vs-aof</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Wed, 24 Jun 2026 14:02:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yGJs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yGJs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!yGJs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!yGJs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!yGJs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yGJs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yGJs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!yGJs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!yGJs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!yGJs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F797aa74e-159f-4c75-b086-4bcfa4cbd175_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The page said the cache lost the last minute of writes. Not the whole dataset, not corruption, just the last minute. We'd been running with the default save points and <code>appendonly</code> off, and when the box got OOM-killed the most recent snapshot was 54 seconds old. Everything written in those 54 seconds was gone. The part that stung is that we'd ticked the "persistence enabled" box in the Helm chart eighteen months earlier and never looked again. Persistence was on. Durability was not. Those are different words, and Redis lets you confuse them for as long as nothing crashes.</p><h2>Two mechanisms, two different lies</h2><p>Redis gives you two on-disk formats and they fail in opposite ways. RDB is a point-in-time snapshot of the whole dataset, written periodically. AOF is an append-only log of every write command, replayed on restart. Most teams pick one without reading what each one promises, and the gap between the promise and the default is where the data goes.</p><p>The same applies to Valkey, byte for byte. Valkey is the BSD-licensed fork of Redis 7.2.4 that AWS, Google, and the Linux Foundation stood up after the 2024 license change, and it inherited the persistence engine unchanged. RDB and AOF files are compatible across both. Everything below is the same engine talking, so when I say Redis, read Valkey too.</p><p>Neither format is the thing people think they bought. RDB loses everything since the last snapshot. AOF in its default config loses up to a second. And both can stall your tail latency hard enough that the SRE on call assumes the node is dying. Which failure mode you hit comes down entirely to the knobs you left at default, so the two formats are worth taking one at a time.</p><h2>RDB: the snapshot that forks your whole heap</h2><p>RDB works by <code>fork()</code>. When a save fires, the parent process forks a child, and the child writes the entire dataset to a temp file while the parent keeps serving traffic. The child sees a frozen copy of memory thanks to copy-on-write: parent and child share physical pages, and a page only gets duplicated when one of them writes to it.</p><p>That copy-on-write detail is the whole story for memory. On a read-mostly instance the fork is nearly free, because almost no pages change while the child writes. But under heavy writes the parent dirties pages fast, the kernel copies each dirtied page, and your resident memory climbs toward 2x the dataset. I've watched a 30 GB Redis with a busy write path balloon past 50 GB of RSS during a save and trip the cgroup limit. The kernel OOM killer doesn't care that the extra memory is transient. It picks the biggest process and kills it, which is your Redis, mid-snapshot.</p><p>Save points are configured as "after N seconds if at least M keys changed". The classic defaults fire roughly every 15 minutes on light churn and as often as every minute under load. Whatever the interval, the data-loss window is "everything since the last completed snapshot". Lose the process between saves and you lose that window. There's no log of the in-between writes, because RDB doesn't keep one.</p><p>The flip side is recovery. Loading an RDB file is fast: it's a compact binary dump, read sequentially, deserialized straight into the keyspace. A multi-gigabyte dataset comes back in seconds. If your priority is "get the cache warm again quickly after a restart" and you can tolerate losing a chunk of recent writes, RDB alone is defensible. Most people who reach for it haven't actually decided they can tolerate that loss, though. They just left it on.</p><h2>AOF: the log that still loses a second</h2><p>AOF logs every write command to a file as it happens. On restart Redis replays the log and rebuilds the exact state. That sounds like real durability, and it can be, but the durability lives entirely in one setting: <code>appendfsync</code>.</p><p>There are three values. With <code>always</code>, Redis calls fsync after every write, so a crash loses at most one command. It's also brutally slow, because you're paying a disk sync on the hot path of every single write, and throughput drops to a fraction of what the same box does otherwise. With <code>no</code>, Redis never explicitly fsyncs and leaves it to the OS, which on Linux means up to 30 seconds of buffered writes can vanish. The middle option, <code>everysec</code>, is the default and the one almost everyone runs: a background thread fsyncs once a second.</p><p><code>everysec</code> is the setting people mean when they say "we have AOF, we're durable". What they actually have is a roughly one-second loss window. Crash at the wrong moment and the writes from the last second that hadn't been synced yet are gone. For a cache or a job queue that's usually fine. The trouble is nobody decided it was fine, they just inherited it, the same way we inherited those RDB save points.</p><p>AOF grows without bound, so Redis periodically rewrites it: it forks a child (yes, the same fork cost as RDB) that writes a compact version representing current state, then the parent appends any commands that arrived during the rewrite. So AOF doesn't escape the fork problem. A write-heavy instance pays the copy-on-write memory tax during every rewrite, not just during snapshots.</p><p>Recovery is where AOF hurts. Replaying a large AOF means re-executing millions of commands one at a time, in order, through the same command parser that handles live traffic. I've seen an AOF-only instance take several minutes to come back where an equivalent RDB loaded in seconds. If your dataset is large and your uptime SLO is tight, that replay time is a real operational cost, not a footnote. It also means a corrupt AOF can stall startup entirely: hit a truncated or malformed entry and Redis refuses to load until you run <code>redis-check-aof --fix</code>, which is exactly the tool you don't want to be learning about during an outage. The file that promised better durability is also the file with more ways to go wrong on the way back up.</p><h2>The hybrid that papers over both</h2><p>Modern Redis defaults to a hybrid that fixes the worst of each. With <code>aof-use-rdb-preamble</code> on, an AOF rewrite writes an RDB-format snapshot as the head of the file, then appends commands in AOF format after it. On restart Redis loads the RDB preamble fast, then replays only the commands logged since the last rewrite.</p><p>You get RDB's quick load and AOF's small loss window in one file. This is the sane default for most workloads, and if you're standing up a new instance it's where I'd start: <code>appendonly yes</code>, <code>appendfsync everysec</code>, RDB preamble on. It doesn't make the fork cost go away, and it doesn't make <code>everysec</code> lose less than a second. It just stops you from having to choose between fast recovery and a bounded loss window. You get both, with the caveats both carry.</p><h2>The fork spike that looks like a dying node</h2><p>The thing that pages people who never touch persistence config is the <code>fork()</code> itself, which is not free even with copy-on-write. The kernel has to copy the parent's page tables, and page-table size scales with how much memory the process maps. On a small instance the fork takes microseconds. On a 100 GB instance it can take hundreds of milliseconds, and during that time the parent process is blocked. Every command in flight waits.</p><p>That shows up in your metrics as a clean p99 jolt: median latency flat, tail latency spiking on a regular cadence that matches your save or rewrite schedule. The first time I saw it I spent an hour chasing GC pauses in the wrong service before noticing the spikes lined up exactly with the snapshot interval. Transparent huge pages make it worse, because each copied page-table entry covers more memory and the COW faults that follow are larger. Disabling THP for Redis is one of the few pieces of folklore that's actually right.</p><p>So persistence isn't just a durability question. The act of persisting is itself a latency event, and on a big instance it's the kind of latency event that wakes people up.</p><h2>Why none of this is durability anyway</h2><p>Step back from the file formats and the harder truth shows up: a single Redis node is never durable, whatever you set <code>appendfsync</code> to. The disk it writes to is the disk that dies. The host it runs on is the host that gets terminated. <code>everysec</code> bounds your loss to a second on a clean crash, but a clean crash isn't the interesting failure. Disk corruption, a yanked EBS volume, a node that vanishes mid-fsync, none of those respect the one-second promise.</p><p>Durability comes from replication, not from a file. You run replicas, writes propagate to them, and the data survives the loss of any one node. But default Redis replication is asynchronous, so a primary can acknowledge a write to your client and die before the replica ever sees it. The client thinks the write succeeded. It didn't survive.</p><p>The <code>WAIT</code> command is the tool that closes that gap. <code>WAIT 1 100</code> blocks until at least one replica has acknowledged the write, or 100 milliseconds pass. It turns the async ack into something closer to a synchronous one, at the cost of the latency you'd expect. Even <code>WAIT</code> isn't a true quorum commit, and it won't save you from a simultaneous loss of primary and replica, but it's the difference between "the write reached one other machine" and "the write reached one machine's page cache". If you genuinely can't lose writes, persistence files are the wrong layer to be arguing about. Replication plus <code>WAIT</code> is the conversation, and even then Redis is a cache that can lose data, not a database that can't.</p><h2>How I actually pick</h2><p>I size the config to what I'd actually lose, not to what feels safe.</p><p>A pure cache that repopulates from a source of truth gets RDB alone, or no persistence at all. The last one of these I set up was a session cache backed by Postgres - if everything in it can be rebuilt, the loss window doesn't matter and the fast restart is what you want.</p><p>A job queue or session store, where a second of loss is annoying but survivable, is where the hybrid earns its keep: <code>appendonly yes</code>, <code>everysec</code>, RDB preamble on. Most workloads I've run land right here, and I've stopped second-guessing it.</p><p>When the data genuinely can't be lost, the file format stops being the question. I push back on that one every time it comes up. You need replicas, <code>WAIT</code> on the critical writes, and an honest look at whether Redis is even the right store.</p><p>And cutting across all of them is the big latency-sensitive instance, where I just watch the fork. Cap memory so a write-burst COW doubling can't OOM you, disable transparent huge pages, and if a single node is too big to fork cheaply, shard it before the page-table copy eats your tail latency.</p><h2>The ones that keep catching people</h2><p>A handful of these show up over and over once you start looking, and I've personally been burned by most.</p><ul><li><p>Believing "persistence enabled" means "durable". It means there's a file on disk that's some amount of stale. The amount is the whole question, and the default answer is "up to a second, or up to your snapshot interval".</p></li><li><p>Then there's running RDB on a write-heavy instance with memory capped at the dataset size. The COW doubling during a save runs you straight into the OOM killer, mid-snapshot, which is the worst possible moment to lose the process.</p></li><li><p>Treating <code>appendfsync everysec</code> as zero loss is the quiet one. It's a one-second window - fine for a cache, quietly wrong for anything you described to a stakeholder as durable.</p></li><li><p>AOF-only with no thought for recovery time bites later: the replay of a large AOF is minutes, not seconds, and you find that out during the incident when you most want the node back.</p></li><li><p>Blame GC or the network for periodic p99 spikes that actually line up with the save schedule, and you'll spelunk in the wrong place for an hour. Overlay the snapshot interval on the latency graph first.</p></li><li><p>Assuming a replica makes you safe with async replication and no <code>WAIT</code>. The primary can ack and die before the replica catches up. The write you were promised is the write you lost.</p></li><li><p>Last, forgetting Valkey behaves identically. The fork is the same fork, the fsync is the same fsync. Switching projects doesn't change the physics, it just changes the license.</p></li></ul><p>Persistence and durability sit one word apart in the docs and a whole incident apart in production. RDB tells you the dataset is safe and means "as of the last snapshot". AOF tells you every write is logged and means "give or take a second". The fork that does both can be the thing that takes your node down. Once you've internalized which lie each format tells, the config stops being a box you tick and starts being a decision you can defend at 3 a.m.</p>]]></content:encoded></item><item><title><![CDATA[Issue #023 - Namespaces aren't free: where 7 TiB of memory was hiding]]></title><description><![CDATA[DaemonSet listwatch, per-node informer cache, Calico netpols, Vector log source, kube-state-metrics cardinality]]></description><link>https://podostack.com/p/issue-023-namespaces-cost-at-scale</link><guid isPermaLink="false">https://podostack.com/p/issue-023-namespaces-cost-at-scale</guid><dc:creator><![CDATA[Ilia Gusev]]></dc:creator><pubDate>Tue, 23 Jun 2026 14:01:37 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fWWp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fWWp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!fWWp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!fWWp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!fWWp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fWWp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png" width="2752" height="1536" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1536,&quot;width&quot;:2752,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fWWp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 424w, https://substackcdn.com/image/fetch/$s_!fWWp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 848w, https://substackcdn.com/image/fetch/$s_!fWWp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!fWWp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fea9bfa12-1ba2-469c-b9e9-f49005f4bc5e_2752x1536.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The capacity dashboard said we had headroom. Forty percent of the cluster's memory free, green across every node, the kind of graph you glance at and stop thinking about. Then a routine DaemonSet rollout - a log-shipper bump, nothing structural - took three nodes into OOMKilled territory inside ninety seconds, and the pods that died weren't the workloads. They were the agents - the log shippers we'd deployed to keep an eye on everything else.</p><p>I spent the first hour convinced it was a leak in the shipper. It wasn't. The agent was doing exactly what it was configured to do, which was keep a local copy of every namespace in the cluster so it could tag each log line with the namespace's labels. We'd grown from a few hundred namespaces to a few thousand over eighteen months, one per tenant, and nobody had drawn the line from "more tenants" to "every DaemonSet pod on every node now holds a few thousand more objects in memory." So the dashboard was right that the RAM was free. It just didn't tell me the agents had quietly reserved most of the rest to mirror a list.</p><p>A namespace looks like the cheapest object in Kubernetes. You make one with a six-line YAML and it costs nothing to create. The bill arrives somewhere else entirely, on a different node, in a process you didn't write, and it scales with a number you stopped tracking.</p><h2>&#128027; Debug Story: reconstructing the 7 TiB that nobody allocated</h2><p>The cleanest public version of this story isn't mine. It's <a href="https://render.com/blog/how-we-found-7-tib-of-memory-just-sitting-around">Render's</a>, and it's worth walking through because they hit the wall at a scale almost nobody else reaches, which makes the mechanism impossible to miss.</p><p>Render runs a PaaS. Every customer service lands in its own namespace, so where a normal cluster lives with tens or low hundreds of namespaces, Render's own framing is that they suspect they're "near the very top" on exactly that axis. When you're at the top of a scaling dimension, the costs that everyone else rounds to zero stop rounding to zero.</p><p>The investigation didn't start with namespaces. It started, as these things do, somewhere adjacent - they were looking at Calico's memory use, because per-namespace NetworkPolicies were the thing visibly stressing it. They built a staging cluster with several hundred thousand namespaces to reproduce, and that test rig is what surfaced the real pattern. The netpols were a scaling factor. They weren't the big one.</p><h3>Following the memory</h3><p>The component that turned out to be holding the memory was <a href="https://vector.dev">Vector</a> - the open-source observability pipeline Render ran as their log collector, deployed as a DaemonSet. That last part is the whole story, so it's worth being precise about why.</p><p>A DaemonSet puts one pod on every node. That's correct and that's the point - a log collector has to be local to the logs. But Vector's <code>kubernetes_logs</code> source was doing something most people never inspect: to enrich each log line with the labels of the namespace it came from, it kept a watch on the namespace API. Not a watch scoped to one namespace. A watch on all of them. And because the collector is a DaemonSet, every node ran its own independent copy of that watch.</p><p>So the cost wasn't per-namespace and it wasn't per-node. It was the product. Each pod held a local cache of every namespace object, and there was one such pod per node. Render's own phrasing is the cleanest summary I've seen: "each of those pods independently performs a listwatch on the same resources. As a result, memory usage increases with the number of nodes." Stack the namespace axis on top of the node axis and you get a memory footprint that grows like nodes times namespaces, in a process whose job description is "ship logs."</p><p>The number that came off the worst pods was around 4 GiB each, just for that cached state. Multiply by the node count on a large cluster and you're into the terabytes on a single cluster before a single log line has been usefully processed.</p><h3>The fix was a feature flag they had to write</h3><p>There was no clever cache-tuning trick here. The namespace watch existed to support one optional behavior - tagging logs with namespace labels - and most of Render's pipeline didn't need it. The fix was to make that watch opt-in instead of always-on. Render contributed the change upstream to Vector: a config option (<code>insert_namespace_fields</code>) that, when off, stops the collector from watching namespaces at all. With the watch gone, there's nothing to cache, so each pod stops keeping its own copy of the cluster's namespace list.</p><p>The catch they hit on the way - and the reason I trust the writeup - is that the first rollout didn't free as much as expected, because they had two <code>kubernetes_logs</code> sources on the user nodes and had only flipped the flag on one. They had to go back and find the second one before the numbers moved, which is the kind of detail a sanitized retrospective drops. Once both were corrected, per-pod memory dropped from those ~4 GiB down to "just a few tens of MiB."</p><p>The aggregate, applied across their whole fleet: one large cluster gave back about 1 TiB, and the total across the infrastructure was roughly 7 TiB of memory that had been sitting there doing nothing but mirroring a list. The 7 TiB is the number that gets quoted, but it's not the part that changed how I look at cluster bills. Not one byte of it showed up under a namespace. It showed up under a log collector, on every node, attributed to a workload that had nothing to do with the tenants whose existence was driving the cost.</p><h3>Links</h3><ul><li><p><a href="https://render.com/blog/how-we-found-7-tib-of-memory-just-sitting-around">How we found 7 TiB of memory just sitting around (Render)</a></p></li><li><p><a href="https://render.com/blog/kubernetes-informers">Kubernetes informers are so easy... to misuse (Render)</a></p></li><li><p><a href="https://vector.dev/docs/reference/configuration/sources/kubernetes_logs/">Vector kubernetes_logs source reference</a></p></li></ul><h2>&#127959;&#65039; Architectural Pattern: why per-namespace cost is roughly O(namespaces)</h2><p>Render's 7 TiB is the extreme case, and the extreme case is useful precisely because it makes a quiet linear cost loud. The same slope exists in your cluster at a few hundred namespaces. It's just under the noise floor, which is exactly why it's dangerous - it grows without ever tripping an alert, until the day it does.</p><p>I want to be careful here, because this topic attracts hand-waving and the readers who'd catch it are the ones I'm writing for. When I went back through the incident, two things had been tangled together that scale very differently with namespace count.</p><h3>The thing that does NOT cost per namespace</h3><p>A DaemonSet pod is per node, not per namespace. This trips people up constantly, so it's worth nailing down: adding a namespace does not add a <code>node-exporter</code> pod, or a <code>kube-proxy</code>, or a CSI node plugin. Those are placed once per node and stay there no matter how many namespaces you create. The pod count of a DaemonSet is a function of your node count, full stop.</p><p>What scales with namespaces isn't the number of agent pods. It's what each of those pods has to know. A DaemonSet that watches a namespace-scoped resource type - namespaces themselves, NetworkPolicies, Services, EndpointSlices - holds a cache whose size tracks the count of those objects, and most of those object types grow as you add namespaces. So the number of agents stays flat while the memory inside each one climbs with every tenant you add. I'd spent the first hour of that incident looking at the flat number, which is exactly where the cost wasn't.</p><h3>Where the real slope lives</h3><p>The mechanism underneath all of this is the client-go informer, and it's worth understanding because it's the same machine inside almost every controller, operator, and agent you run. An informer opens a watch against the API server and keeps a local in-memory cache of every object matching its scope, kept current by the stream of ADDED/MODIFIED/DELETED events. The cache is the feature - it's what lets a controller answer "what Services exist" without hammering the API server. But by default that cache is cluster-wide. Every object of the watched type, in every namespace, materialized in the process's heap, deep-copied out of the wire format into Go structs.</p><p>So the cost of one informer is roughly: number of objects in scope, times the size of each object, times the number of processes running that informer. Add a namespace and you've added its default ServiceAccount, its EndpointSlices, whatever NetworkPolicies and Services and ConfigMaps come with a tenant. Every informer cluster-watching those types just got a little bigger, in every replica, on every node if the watcher is a DaemonSet. None of it is a leak. It's the cache doing its job at a scale the cache's author didn't picture.</p><p>There's a second cost that doesn't show up as agent memory at all, and it's the one that bites the control plane: the API server has to serve all those watches. Each watching client holds an open connection and gets a stream of every relevant change. The kube-apiserver maintains its own watch cache to fan this out, but the connection count and the event volume both climb with watchers times objects. The failure mode Render flagged is the rollout: when a DaemonSet's pods all restart at once, they re-LIST the world simultaneously, and that thundering herd of full LISTs can, in their words, "overwhelm the server with requests and cause real disruption." The steady state is expensive. The synchronized cold start is the part that takes the cluster down.</p><h3>The things that genuinely track namespace count</h3><p>Walk the per-namespace machinery and the linear costs stack up, none of them individually alarming:</p><p>When you create a namespace, the controller manager's serviceaccount-controller and token-controller spin up a <code>default</code> ServiceAccount and its token plumbing for it - automatic, per namespace, forever. CoreDNS watches Services and EndpointSlices cluster-wide; <a href="https://coredns.io/2018/11/15/scaling-coredns-in-kubernetes-clusters/">its own sizing guidance</a> is <code>(Pods + Services) / 1000 + 54</code> MB, so more tenants means more Services means a fatter DNS cache on every CoreDNS replica. kube-state-metrics builds a time series per object across every namespace, so its scrape payload and memory climb with the object count until you're sharding it across a StatefulSet just to keep the <code>/metrics</code> endpoint from timing out. Any operator that watches all namespaces - cert-manager, an ingress controller, a service-mesh control plane, a monitoring agent - carries its own informer cache with the same slope. And admission webhooks configured cluster-wide get invoked for matching operations in every namespace, so a webhook that was cheap at fifty tenants is doing real work at five thousand.</p><p>No single line of that is a problem. The architecture problem is that they're all the same slope at once, and a namespace is the unit that pays into every one of them simultaneously while appearing, on its own ledger, to cost nothing. You're not provisioning a namespace. You're provisioning a row in a dozen different caches you don't own and a dozen different watch streams you can't see.</p><h3>Links</h3><ul><li><p><a href="https://pkg.go.dev/k8s.io/client-go/tools/cache">client-go: tools/cache reflector and informer internals</a></p></li><li><p><a href="https://kube.rs/controllers/optimization/">controller-runtime cache optimization (kube.rs)</a></p></li><li><p><a href="https://github.com/coredns/deployment/blob/master/kubernetes/Scaling_CoreDNS.md">Scaling CoreDNS and the memory formula</a></p></li><li><p><a href="https://coredns.io/2018/11/15/scaling-coredns-in-kubernetes-clusters/">Scaling CoreDNS in Kubernetes Clusters (coredns.io)</a></p></li><li><p><a href="https://aws.github.io/aws-eks-best-practices/scalability/docs/cluster-services/">EKS best practices: cluster services scalability</a></p></li><li><p><a href="https://kubernetes.io/docs/reference/access-authn-authz/service-accounts-admin/">Managing ServiceAccounts (per-namespace default)</a></p></li></ul><h2>&#128736;&#65039; The practical part: measuring the tax, then cutting it</h2><p>The reason this cost stays invisible is that nothing attributes it to namespaces. So before any mitigation, the first job is to make the slope visible, and most of that you can do with tools already in the cluster.</p><h3>Finding the slope</h3><p>The first thing I checked that night was a question I'd never bothered to ask before the cluster forced it: how many namespaces do we actually have, and is that number growing.</p><pre><code>$ kubectl get ns --no-headers | wc -l
3847

$ kubectl get networkpolicies -A --no-headers | wc -l
3812

$ kubectl get serviceaccounts -A --no-headers | wc -l
9216</code></pre><p>If those counts are in the thousands and climbing, the slope is already in your cluster whether or not it's hurting yet. The next question is which processes are paying for it. The agents to suspect are the ones that are both cluster-watching and replicated per node - DaemonSets - so sort their pods by memory:</p><pre><code>$ kubectl top pods -A --sort-by=memory | grep -E 'vector|fluent|cilium|calico|otel|datadog'
kube-system   cilium-9x4kp       412Mi
logging       vector-agent-7fh2k 3964Mi
logging       vector-agent-2bnq9 3871Mi</code></pre><p>A log agent sitting at 4 GiB while the workloads it's collecting from use a fraction of that is the signature. When you see one DaemonSet pod an order of magnitude heavier than its peers across the fleet, and the number tracks your namespace count rather than your log volume, you've found a process holding a mirror of the cluster it doesn't need. For the deeper attribution Render did - confirming which watch is responsible - a continuous profiler like Pyroscope or a <code>go tool pprof</code> heap dump against the agent's debug endpoint shows the informer cache as a fat allocation, but <code>kubectl top</code> plus the namespace count usually tells you where to look first.</p><h3>The fixes, in order of how much you have to give up</h3><p>The cheapest fix is the one Render shipped: stop the watch you don't need. If an agent is caching a resource type only to support a feature you're not using, turn the feature off. The namespace-label enrichment in a log collector is the classic example - the watch that powered it was free until namespaces stopped being free.</p><p>When you do need the watch, scope it. An informer doesn't have to be cluster-wide. client-go and controller-runtime both let you constrain a watch with a label selector, a field selector, or a fixed list of namespaces, and the cache then only holds what matches. A label selector is the most expressive lever - field selectors can only filter a handful of indexed fields - so the durable pattern is to label the objects an agent actually cares about and watch by that label:</p><pre><code># controller-runtime: cache only the namespaces a tenant-agent owns,
# instead of mirroring all 3,847 of them
cache:
  byObject:
    namespaces:
      label: "podostack.com/tenant-agent in (true)"</code></pre><p>That single constraint is the difference between a cache that grows with the whole cluster and one that grows with the slice you operate. It's also the change with the best ratio - you keep the functionality, you drop the slope.</p><p>Past tuning, the lever is structural: stop treating the namespace as the tenant boundary. A namespace was never an isolation primitive; it's a name scope with some RBAC and quota bolted on, and every cluster-wide controller still sees straight through it. If the reason you have thousands of namespaces is hard multi-tenancy, the honest tool is a virtual cluster - vCluster gives each tenant its own API server and control plane, so a tenant's Services and NetworkPolicies live in that tenant's control plane and stop landing in every host-level informer's cache. The cost is real and worth naming: you're now debugging two layers, and your host monitoring won't see into the guest clusters without extra wiring. Hierarchical namespaces (HNC) sit between the two - they organize namespaces into a tree with inherited policy, which helps governance but does nothing for the per-namespace informer cost, because to every cluster-wide watcher an HNC namespace is just another namespace. Worth knowing what each one actually buys, because only one of them touches the slope this issue is about.</p><p>The decision compresses to this. A few hundred namespaces and flat: do nothing, the cost is real but under your noise floor. Thousands and climbing, with soft multi-tenancy: audit your DaemonSets and operators, scope or kill the cluster-wide watches, and you'll likely find your own smaller version of the 7 TiB. Thousands and climbing because you're running a hard-multi-tenant platform: the namespace was the wrong boundary, and a tenancy layer that gives each tenant a real control plane is what stops the slope at the source rather than tuning it down one informer at a time.</p><h3>Links</h3><ul><li><p><a href="https://www.vcluster.com/blog/comparing-multi-tenancy-options-in-kubernetes">vCluster: comparing multi-tenancy options</a></p></li><li><p><a href="https://github.com/kubernetes-sigs/controller-runtime/issues/778">controller-runtime issue #778: watch with a label selector</a></p></li><li><p><a href="https://github.com/kubernetes-sigs/hierarchical-namespaces">Hierarchical Namespace Controller (HNC)</a></p></li><li><p><a href="https://github.com/kubernetes/kube-state-metrics#horizontal-sharding">kube-state-metrics sharding for large clusters</a></p></li></ul><h2>Until next week</h2><p>The part I keep coming back to: the cost was never where the abstraction said it was. You provision a namespace, the bill lands on a log collector on a node you weren't looking at, and the only way to find it is to stop trusting the org chart of your cluster and follow the memory instead. Go run <code>kubectl get ns | wc -l</code> on your busiest cluster and then <code>kubectl top pods -A --sort-by=memory</code>. If those two numbers are correlated, you've got a smaller version of this sitting in your fleet right now.</p><p>Next Tuesday we stay with things that aren't supposed to be visible, and look at how eBPF reads TLS traffic without ever touching the private key - decrypting at the syscall boundary, before the bytes hit the network. See you then.</p><p>- Ilia</p>]]></content:encoded></item></channel></rss>