<?xml version="1.0" encoding="UTF-8"?>
    <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/">
      <channel>
        <title>SRE and DevOps notes from production | Amjed Allaya</title>
        <link>https://amjedallaya.com/blog</link>
        <description>Evidence-backed production field notes about incident response, observability, safe automation, Java performance, cloud migrations, and Kubernetes.</description>
        <language>en</language>
        <copyright>© 2026 Amjed Allaya</copyright>
        <generator>Amjed Allaya portfolio content pipeline</generator>
        <ttl>60</ttl>
        <lastBuildDate>Sat, 19 Sep 2026 00:00:00 GMT</lastBuildDate>
        <atom:link href="https://amjedallaya.com/feed.xml" rel="self" type="application/rss+xml" />
        
      <item>
        <title>Top 5 terminal emulators in 2026: an honest comparison</title>
        <link>https://amjedallaya.com/blog/best-terminal-emulators-2026</link>
        <guid isPermaLink="true">https://amjedallaya.com/blog/best-terminal-emulators-2026</guid>
        <pubDate>Sat, 19 Sep 2026 00:00:00 GMT</pubDate>
        <atom:updated>2026-09-19T00:00:00Z</atom:updated>
        <dc:creator>Amjed Allaya</dc:creator>
        <description>A practical 2026 comparison of Ghostty, WezTerm, kitty, Warp, and Automexia, focused on strengths, limitations, and the workflow each terminal suits best.</description>
        <category>Developer tooling</category><category>Terminal emulators</category><category>Automexia</category><category>Ghostty</category><category>WezTerm</category><category>kitty</category><category>Warp</category>
        <media:content url="https://amjedallaya.com/images/blog/terminal-emulators-2026-cover.webp" type="image/webp" medium="image" width="1672" height="941">
          <media:title>Top 5 terminal emulators in 2026: an honest comparison</media:title>
          <media:description type="plain">An engineer compares connected terminal windows displaying the Ghostty, Automexia, WezTerm, kitty, and Warp logos.</media:description>
        </media:content>
        <content:encoded><![CDATA[<p><small>Published <time datetime="2026-09-19">September 19, 2026</time> · Updated <time datetime="2026-09-19">September 19, 2026</time></small></p><figure><img src="https://amjedallaya.com/images/blog/terminal-emulators-2026-cover.webp" alt="An engineer compares connected terminal windows displaying the Ghostty, Automexia, WezTerm, kitty, and Warp logos." width="1672" height="941"><figcaption>The best terminal is the one whose operating model matches the way you actually work.</figcaption></figure>
<style>
.terminal-comparison-article {
  font-size: 1.0625rem;
  line-height: 1.78;
}

.terminal-comparison-article p {
  margin-bottom: 1.15rem;
}

.terminal-comparison-article h2 {
  line-height: 1.2;
}

.terminal-comparison-article h3 {
  font-size: 1.14rem;
  line-height: 1.35;
  margin: 0 0 0.75rem;
}

.terminal-review {
  display: grid;
  grid-template-columns: minmax(0, 1fr) minmax(260px, 36%);
  gap: 2.25rem;
  align-items: start;
  margin: 0.75rem 0 1.25rem;
}

.terminal-review__lead {
  font-size: 1.08em;
  line-height: 1.75;
}

.terminal-review__tagline {
  font-size: 1.08em;
  margin-bottom: 0.8rem;
}

.terminal-review__columns {
  display: grid;
  grid-template-columns: repeat(2, minmax(0, 1fr));
  gap: 1.75rem;
  margin-top: 1.5rem;
}

.terminal-review__columns ul {
  margin: 0;
  padding-left: 1.15rem;
}

.terminal-review__columns li {
  margin-bottom: 0.58rem;
}

.terminal-review__preview {
  margin: 0;
}

.terminal-review__preview img {
  display: block;
  width: 100%;
  height: auto;
  border-radius: 14px;
}

.terminal-review__preview figcaption {
  font-size: 0.9em;
  line-height: 1.5;
  margin-top: 0.65rem;
  opacity: 0.78;
}

.terminal-review__action {
  margin-top: 1.35rem;
  font-weight: 600;
}

.terminal-comparison-note {
  margin-top: 1.5rem;
}

@media (max-width: 820px) {
  .terminal-comparison-article {
    font-size: 1rem;
  }

  .terminal-review {
    grid-template-columns: 1fr;
    gap: 1.35rem;
  }

  .terminal-review__columns {
    grid-template-columns: 1fr;
    gap: 1rem;
  }
}
</style>
<div class="terminal-comparison-article"><div><p>A terminal is something you may keep open for hours, so the right choice is less about a giant feature list and more about how comfortably it fits your daily workflow.</p><p>This comparison looks at Ghostty, WezTerm, kitty, Warp, and Automexia. They all handle the core terminal job, but they make different choices around native desktop behavior, configuration, automation, connected features, and session organization.</p><p>Use the quick comparison to build a shortlist. Then jump directly to the individual terminal sections that match your platform and the way you work.</p></div><section id="overview"><h2>Overview</h2><p>The comparison is deliberately practical. I look at four things: whether the terminal supports the operating systems you use, how much configuration it asks you to maintain, whether its workflow stays close to a traditional shell, and what additional account or network boundaries it introduces.</p><blockquote>Disclosure: I am the creator of Automexia. Its section includes firsthand product context. The other products are assessed from their current official documentation, and Automexia&#x27;s Public Alpha status is treated as a real limitation throughout this article.</blockquote><p>I checked the public release state on September 19, 2026. Planned features do not count as shipped. I also avoid a synthetic “fastest terminal” ranking because rendering results depend on hardware, fonts, scrollback, and workload.</p><p><strong>Sources: </strong><a href="#reference-automexia-docs">Automexia</a>, <a href="#reference-automexia-repository">GitHub</a>, <a href="#reference-ghostty-about">Ghostty</a>, <a href="#reference-wezterm-features">WezTerm</a>, <a href="#reference-kitty-overview">kitty</a>, <a href="#reference-warp-getting-started">Warp</a></p></section><section id="quick-comparison"><h2>Quick comparison</h2><p>Start here. The purpose of this table is not to pick a universal winner; it is to remove options that clearly do not fit your platform or workflow before you spend time on the detailed reviews.</p><figure><figcaption><strong>Terminal comparison at a glance</strong><p>A compact shortlist before the detailed reviews.</p></figcaption><table><thead><tr><th scope="col">Terminal</th><th scope="col">Best for</th><th scope="col">Main limitation</th><th scope="col">Platforms</th></tr></thead><tbody><tr><th scope="row">Ghostty</th><td>Native-feeling desktop behavior with focused defaults</td><td>No Windows app; Linux packaging depends on the distribution</td><td>macOS, Linux</td></tr><tr><th scope="row">WezTerm</th><td>A programmable setup shared across operating systems</td><td>Lua configuration can become a project of its own</td><td>Windows, macOS, Linux, BSD</td></tr><tr><th scope="row">kitty</th><td>Terminal-native automation, protocols, and extensions</td><td>No native Windows app and a real learning curve</td><td>Linux, macOS, BSD</td></tr><tr><th scope="row">Warp</th><td>A terminal workspace with an editor, command blocks, and agents</td><td>Connected features widen the network and data boundary</td><td>Windows, macOS, Linux</td></tr><tr><th scope="row">Automexia</th><td>Organized local terminal sessions with the shell remaining in control</td><td>Public Alpha; stable Windows and macOS releases are not published</td><td>Linux x64, Linux Arm64</td></tr></tbody></table></figure><p class="terminal-comparison-note">The large stacked product-card list has been removed from the beginning. Each visual now appears only inside the relevant product section, where it supports the review instead of competing with the rest of the page.</p></section><section id="ghostty"><h2>1. Ghostty</h2><div class="terminal-review"><div><p class="terminal-review__tagline"><strong>Fast, native-feeling, and focused.</strong></p><p class="terminal-review__lead"><a href="https://ghostty.org/">Ghostty</a> is designed to feel like part of the desktop around it rather than a separate cross-platform interface. It combines a shared Zig core with AppKit and SwiftUI on macOS and GTK4 on Linux, so tabs, splits, shortcuts, and window behavior follow the host operating system.</p><div class="terminal-review__columns"><div><h3>Highlights</h3><ul><li>Focused defaults with native tabs and splits.</li><li>Hyperlinks, themes, and graphics support are available without a large configuration first.</li><li>Its project description keeps performance claims bounded rather than treating one benchmark as universal.</li></ul></div><div><h3>Things to consider</h3><ul><li>There is no Windows app.</li><li>On Linux, package availability and update speed depend on the distribution.</li><li>The macOS and Linux interfaces are intentionally different.</li></ul></div></div><p class="terminal-review__action"><a href="https://ghostty.org/">Visit Ghostty →</a></p></div><figure class="terminal-review__preview"><img src="/images/blog/terminal-ghostty-preview.webp" alt="Official Ghostty artwork showing a ghost silhouette drawn with terminal characters." width="1440" height="810" loading="lazy"/><figcaption>Best fit: developers on macOS or Linux who value native desktop behavior and focused defaults.</figcaption></figure></div><p><strong>Sources: </strong><a href="#reference-ghostty-about">Ghostty</a>, <a href="#reference-ghostty-install">Ghostty</a></p></section><section id="wezterm"><h2>2. WezTerm</h2><div class="terminal-review"><div><p class="terminal-review__tagline"><strong>Programmable, flexible, and genuinely cross-platform.</strong></p><p class="terminal-review__lead"><a href="https://wezterm.org/">WezTerm</a> is the clearest fit here if you want one terminal setup to travel across several operating systems. Its documented platform range includes Windows, macOS, Linux, FreeBSD, and NetBSD, and it combines terminal emulation with local and remote multiplexing.</p><div class="terminal-review__columns"><div><h3>Highlights</h3><ul><li>Lua configuration can describe key bindings, appearance, events, launch menus, and platform conditions.</li><li>SSH domains, images, ligatures, search, and copy modes live in the same programmable environment.</li><li>Windows users get ConPTY support and official packages.</li></ul></div><div><h3>Things to consider</h3><ul><li>A powerful Lua config can grow into a small software project.</li><li>Copied snippets and event handlers still need long-term ownership.</li><li>Remote multiplexing deserves a real reconnection and recovery test before incident use.</li></ul></div></div><p class="terminal-review__action"><a href="https://wezterm.org/">Visit WezTerm →</a></p></div><figure class="terminal-review__preview"><img src="/images/blog/terminal-wezterm-preview.webp" alt="WezTerm on macOS displaying Rust source code in Vim." width="1440" height="810" loading="lazy"/><figcaption>Best fit: developers who value a programmable configuration shared across multiple operating systems.</figcaption></figure></div><p><strong>Sources: </strong><a href="#reference-wezterm-features">WezTerm</a>, <a href="#reference-wezterm-install">WezTerm</a></p></section><section id="kitty"><h2>3. kitty</h2><div class="terminal-review"><div><p class="terminal-review__tagline"><strong>Terminal-native power without turning the terminal into a separate workspace.</strong></p><p class="terminal-review__lead"><a href="https://sw.kovidgoyal.net/kitty/">kitty</a> stays close to traditional terminal workflows while adding GPU rendering, tabs, layouts, sessions, remote control, and small tools called kittens.</p><div class="terminal-review__columns"><div><h3>Highlights</h3><ul><li>Its graphics and keyboard protocols have influenced other terminal tools.</li><li>Kittens add focused capabilities such as diffing, Unicode input, themes, and remote-file handling.</li><li>Much of the experience remains accessible through configuration files and scripts.</li></ul></div><div><h3>Things to consider</h3><ul><li>There is no native Windows app.</li><li>Its vocabulary and deeper features take time to learn.</li><li>Remote control should be configured deliberately because it expands what another process can ask the terminal to do.</li></ul></div></div><p class="terminal-review__action"><a href="https://sw.kovidgoyal.net/kitty/">Visit kitty →</a></p></div><figure class="terminal-review__preview"><img src="/images/blog/terminal-kitty-preview.webp" alt="kitty&#x27;s Unicode input tool searching for cat emoji inside the terminal." width="1440" height="810" loading="lazy"/><figcaption>Best fit: terminal-first users who actively use scripting, protocols, and extensibility.</figcaption></figure></div><p><strong>Sources: </strong><a href="#reference-kitty-overview">kitty</a>, <a href="#reference-kitty-install">kitty</a></p></section><section id="warp"><h2>4. Warp</h2><div class="terminal-review"><div><p class="terminal-review__tagline"><strong>A broader development workspace built around the terminal.</strong></p><p class="terminal-review__lead"><a href="https://www.warp.dev/">Warp</a> goes further than the other terminals in this comparison by combining the terminal with an editor, file tree, structured command blocks, and agent workflows.</p><div class="terminal-review__columns"><div><h3>Highlights</h3><ul><li>Command output stays attached to the command that produced it.</li><li>The editor and agent workflow can reduce context switching during code changes and review.</li><li>The core terminal supports familiar shells, and offline use remains available after the first launch.</li></ul></div><div><h3>Things to consider</h3><ul><li>AI, sharing, and collaboration widen the network and data boundary.</li><li>Teams still need to compare Warp&#x27;s privacy and telemetry controls with their own policy.</li><li>Its workflow is intentionally broader than a traditional terminal-only model.</li></ul></div></div><p class="terminal-review__action"><a href="https://www.warp.dev/">Visit Warp →</a></p></div><figure class="terminal-review__preview"><img src="/images/blog/terminal-warp-preview.webp" alt="Warp workspace showing a project tree and structured command results." width="1440" height="810" loading="lazy"/><figcaption>Best fit: developers who specifically want editor, command-block, and agent features inside the terminal workspace.</figcaption></figure></div><p><strong>Sources: </strong><a href="#reference-warp-getting-started">Warp</a>, <a href="#reference-warp-install">Warp</a>, <a href="#reference-warp-offline">Warp</a>, <a href="#reference-warp-privacy">Warp</a></p></section><section id="automexia"><h2>5. Automexia</h2><div class="terminal-review"><div><p class="terminal-review__tagline"><strong>Organized local terminal work with the shell remaining in control.</strong></p><p class="terminal-review__lead"><a href="https://www.automexia.com/docs">Automexia</a> focuses on keeping related local terminal sessions organized. Windows, global tabs, split panes, and pane-local tabs give native shell sessions stable places without replacing the shell with a separate command model.</p><div class="terminal-review__columns"><div><h3>Highlights</h3><ul><li>Search, scrollback, images, and configuration remain local.</li><li>No account, hosted service, provider login, or model is required for the released Linux application.</li><li>Terminal output is treated as untrusted display data, so links, previews, and search results cannot silently become shell input.</li></ul></div><div><h3>Things to consider</h3><ul><li>Automexia v0.4.0 is Public Alpha.</li><li>Published packages support Linux x64 and Arm64; stable Windows and macOS installers are not available.</li><li>The community and third-party compatibility evidence are smaller than the mature alternatives in this comparison.</li></ul></div></div><p class="terminal-review__action"><a href="https://www.automexia.com/docs">Visit Automexia →</a></p></div><figure class="terminal-review__preview"><img src="/images/blog/terminal-automexia-preview.webp" alt="Automexia preview showing a dark terminal workspace beside the product name." width="1440" height="810" loading="lazy"/><figcaption>Best fit: Linux users who want organized local sessions and are comfortable evaluating a Public Alpha release.</figcaption></figure></div><p><strong>Sources: </strong><a href="#reference-automexia-docs">Automexia</a>, <a href="#reference-automexia-features">Automexia</a>, <a href="#reference-automexia-repository">GitHub</a>, <a href="#reference-automexia-platforms">GitHub</a></p></section><section id="how-to-choose"><h2>How to choose</h2><p>A useful comparison should leave you with a shortlist, not a universal winner. Start by removing anything that fails a hard platform or security requirement. Then compare the workflow you actually want with the trade-off you would need to accept.</p><figure><figcaption><strong>Start with your workflow</strong><p>Use the first column to find the closest match, then verify the trade-off before switching.</p></figcaption><table><thead><tr><th scope="col">If your priority is...</th><th scope="col">Look closely at...</th><th scope="col">Verify before switching...</th></tr></thead><tbody><tr><th scope="row">Native desktop behavior on macOS or Linux</th><td>Ghostty</td><td>Whether the missing Windows support or Linux packaging model affects you</td></tr><tr><th scope="row">One programmable setup across many systems</th><td>WezTerm</td><td>Whether you want to maintain a Lua-based configuration long term</td></tr><tr><th scope="row">Terminal-native extensions and automation</th><td>kitty</td><td>Whether its learning curve and lack of native Windows support are acceptable</td></tr><tr><th scope="row">A terminal workspace with editor and agent features</th><td>Warp</td><td>Whether its broader connectivity and data boundary suit your environment</td></tr><tr><th scope="row">Organized local Linux sessions with the shell still in control</th><td>Automexia</td><td>Whether a Public Alpha release is appropriate for your current needs</td></tr></tbody></table></figure></section><section id="evaluation-test"><h2>Before you switch</h2><p>Do not replace your current terminal because a comparison table looks convincing. Give each finalist one real hour with the shell and layout you already use.</p><figure><figcaption><strong>A one-hour terminal evaluation</strong><p>Run the same small test for every candidate.</p></figcaption><table><thead><tr><th scope="col">Try</th><th scope="col">Good sign</th><th scope="col">Stop if</th></tr></thead><tbody><tr><th scope="row">Your normal shell</th><td>Aliases, quoting, and paste work as expected</td><td>Basic habits become ambiguous</td></tr><tr><th scope="row">A deployment-style layout</th><td>The active session stays obvious</td><td>Context becomes harder to follow</td></tr><tr><th scope="row">A network or configuration failure</th><td>The failure is clear and recovery is quick</td><td>The recovery path is unclear</td></tr></tbody></table></figure><p>If the test goes well, keep the old and new terminals side by side for a few days. The best long-term choice is the one that becomes predictable enough that you stop thinking about the terminal and get back to the work inside it.</p></section></div><section><h2>Key takeaways</h2><ul><li>Ghostty is a strong fit for developers who want a native-feeling macOS or Linux terminal.</li><li>WezTerm is a strong fit for developers who want one programmable setup across several operating systems.</li><li>kitty suits terminal-first users who value protocols, automation, and scriptable extensions.</li><li>Warp makes the most sense when the editor, command blocks, and agent workflow are part of the reason to switch.</li><li>Automexia focuses on organized local Linux sessions, but its Public Alpha status is an important practical limit.</li><li>The right choice is the terminal whose trade-offs fit your real workflow, not the one with the longest feature list.</li></ul></section><section><h2>References &amp; revisions</h2><ol><li id="reference-automexia-docs"><a href="https://www.automexia.com/docs">Automexia Terminal documentation</a> — Automexia</li><li id="reference-automexia-features"><a href="https://www.automexia.com/docs/features">Automexia feature catalog</a> — Automexia</li><li id="reference-automexia-repository"><a href="https://github.com/AmjedAllaya/automexia">Automexia Terminal source repository</a> — GitHub</li><li id="reference-automexia-platforms"><a href="https://github.com/AmjedAllaya/automexia/blob/main/docs/PLATFORMS.md">Automexia platform support</a> — GitHub</li><li id="reference-ghostty-about"><a href="https://ghostty.org/docs/about">About Ghostty</a> — Ghostty</li><li id="reference-ghostty-install"><a href="https://ghostty.org/docs/install/binary">Ghostty binaries and packages</a> — Ghostty</li><li id="reference-wezterm-features"><a href="https://wezterm.org/features.html">WezTerm features</a> — WezTerm</li><li id="reference-wezterm-install"><a href="https://wezterm.org/installation.html">WezTerm downloads and supported platforms</a> — WezTerm</li><li id="reference-kitty-overview"><a href="https://sw.kovidgoyal.net/kitty/overview/">kitty overview</a> — kitty</li><li id="reference-kitty-install"><a href="https://sw.kovidgoyal.net/kitty/binary/">Install kitty</a> — kitty</li><li id="reference-warp-getting-started"><a href="https://docs.warp.dev/">Getting started with Warp</a> — Warp</li><li id="reference-warp-install"><a href="https://docs.warp.dev/getting-started/quickstart/installation-and-setup/">Warp installation and setup</a> — Warp</li><li id="reference-warp-offline"><a href="https://docs.warp.dev/support-and-community/troubleshooting-and-support/using-warp-offline">Using Warp offline</a> — Warp</li><li id="reference-warp-privacy"><a href="https://docs.warp.dev/support-and-community/privacy-and-security/privacy">Warp privacy and telemetry controls</a> — Warp</li></ol><h3>Revision history</h3><p><time datetime="2026-09-19">2026-09-19</time> — Simplified the comparison layout, kept the existing branding direction, increased reading size and spacing, replaced the stacked product-card list with a compact comparison table, and gave each terminal its own spacious review section.</p></section>]]></content:encoded>
      </item>
      <item>
        <title>Incident response across a multi-cloud migration</title>
        <link>https://amjedallaya.com/blog/calm-multi-cloud-incident-response</link>
        <guid isPermaLink="true">https://amjedallaya.com/blog/calm-multi-cloud-incident-response</guid>
        <pubDate>Tue, 18 Aug 2026 00:00:00 GMT</pubDate>
        <atom:updated>2026-08-25T00:00:00Z</atom:updated>
        <dc:creator>Amjed Allaya</dc:creator>
        <description>How I find the first broken boundary in a multi-cloud message flow and recover without making the migration harder to reverse.</description>
        <category>Incident response</category><category>Incident command</category><category>Azure migration</category><category>Kafka</category><category>RabbitMQ</category><category>Kubernetes</category>
        <media:content url="https://amjedallaya.com/images/blog/incident-response-cover.webp" type="image/webp" medium="image" width="1672" height="941">
          <media:title>Incident response across a multi-cloud migration</media:title>
          <media:description type="plain">An operator routes traffic around a failed service while preserving the rest of the user journey.</media:description>
        </media:content>
        <content:encoded><![CDATA[<p><small>Published <time datetime="2026-08-18">August 18, 2026</time> · Updated <time datetime="2026-08-25">August 25, 2026</time></small></p><figure><img src="https://amjedallaya.com/images/blog/incident-response-cover.webp" alt="An operator routes traffic around a failed service while preserving the rest of the user journey." width="1672" height="941"><figcaption>A safe response isolates the failing path and preserves a reversible route back.</figcaption></figure><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/incident-multicloud-readiness.webp"/><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/incident-first-fifteen-minutes.webp"/>
<div><p>The hardest part of a multi-cloud migration is keeping the message path clear while services and infrastructure move at different speeds.</p><p>Before moving traffic, I answer three questions. Which path is active? What will prove that a message completed? What exact action returns traffic if the new path fails? Those answers turn a risky migration into a sequence the team can observe and reverse.</p></div>
<section id="production-problem"><h2>The production problem is an ambiguous message path</h2><p>A green health check proves that the connector process is running. It does not prove that the expected message reached the correct consumer. A pod can be ready while its consumer is attached to the wrong topic, so process health and flow health must be checked separately.</p><p>I follow one known message from the producer to the business outcome. An opaque correlation ID travels with the message and lets an operator find it without searching for customer data. At each boundary, I ask whether the message entered and whether it left. The first boundary without evidence becomes the starting point for the investigation. I also record which team owns that boundary, so the incident does not stall while people search for the right operator.</p><p>The decision I changed was where recovery ends. Connector health is still useful, but it is too narrow to be the final checkpoint. I now end the evidence path at the business outcome. This prevents a healthy process from hiding work that is delayed, duplicated, or missing.</p><figure><figcaption><strong>Evidence checkpoints for a migrating flow</strong><p>The control is useful only when it proves progress at the next boundary, not merely local health.</p></figcaption><table><thead><tr><th scope="col">Checkpoint</th><th scope="col">Evidence to capture</th><th scope="col">What it rules out</th></tr></thead><tbody><tr><th scope="row">Producer</th><td>Find the message by its correlation ID and confirm the destination.</td><td>Shows whether the application emitted the message.</td></tr><tr><th scope="row">Broker boundary</th><td>Compare the source position with the replicated position.</td><td>Shows whether the migration path moved the message.</td></tr><tr><th scope="row">Consumer</th><td>Track the known message while confirming that the consumer offset continues to advance. The offset is the broker position saved for that consumer group.</td><td>Shows whether a healthy process is actually consuming.</td></tr><tr><th scope="row">Business outcome</th><td>Confirm the terminal event or locate the item in a reconciliation report.</td><td>Prevents transport success from being mistaken for user success.</td></tr></tbody></table></figure><blockquote>During a migration, “the connector is up” is an infrastructure statement. “The expected business event completed once” is a recovery statement.</blockquote></section>
<section id="decision-record"><h2>Write the cutover contract before touching production</h2><p>The most useful migration document I keep is one page for one flow. The first line names the path that currently owns the traffic and the path that will replace it. The next line names the person who can approve the move and how long the team will observe it.</p><figure><img src="https://amjedallaya.com/images/blog/incident-multicloud-readiness.webp" alt="An operator validates each step before moving a flow between cloud environments." width="1536" height="1024"/><figcaption>A cutover is ready when the new path works and the return path has been tested.</figcaption></figure><p>The record then defines proof. For example, it may require the consumer offset to keep advancing and a business completion event to appear. It also defines one stop signal and one tested route back. The thresholds stay as placeholders until the service owner chooses values that fit the workload.</p><p>This removes a dangerous debate from the incident channel. If the new consumer falls behind before cutover, the old path remains authoritative while the team investigates. If traffic already moved and message correctness becomes uncertain, the team follows the tested return route. A human still makes the decision, but the evidence and safe options are already clear.</p><figure><figcaption><strong>Automexia code terminal — Sanitized cutover decision record</strong> <small>YAML</small></figcaption><pre><code>flow: &lt;business-flow&gt;
current_source_of_truth: on_prem
target: azure
owner: &lt;team-and-on-call&gt;
observe:
  - producer_send_rate
  - mirror_lag
  - consumer_offset_progress
  - business_completion_rate
abort_when:
  mirror_lag: &lt;agreed-workload-limit&gt;
  duplicate_rate: &lt;agreed-workload-limit&gt;
  completion_rate: &lt;agreed-slo-limit&gt;
reverse:
  route: &lt;tested-return-path&gt;
  data_check: &lt;reconciliation-query&gt;</code></pre><p>This is a sanitized template. Every placeholder must be agreed and tested for the specific flow.</p></figure><p><strong>Sources: </strong><a href="#reference-kafka-mirrormaker">Apache Kafka</a></p></section>
<section id="first-fifteen-minutes"><h2>Use the first fifteen minutes to remove ambiguity</h2><p>When a migrated flow degrades, I do not start by blaming a cloud platform or broker. I first name the exact business flow and the direction in which it is failing. Then I record when the problem began and the last outcome we know succeeded. This gives every team the same incident to investigate.</p><figure><img src="https://amjedallaya.com/images/blog/incident-first-fifteen-minutes.webp" alt="An incident commander assigns clear investigation and communication responsibilities." width="1535" height="1024"/><figcaption>Coordination and investigation need separate owners when several platforms are involved.</figcaption></figure><p>The sequence below is a response drill derived from the migration work. It is intentionally evidence-first. Root cause can wait; an incorrect traffic move can make reconciliation much harder.</p><ul><li>Incident commander: sets the next objective and decides whether a production change is allowed.</li><li>Flow lead: follows one message path end to end and prevents unrelated debugging branches.</li><li>Platform owners: inspect their boundary and return timestamped evidence instead of saying that it looks healthy.</li><li>Scribe: records each query and production change with a timestamp. This makes later reconciliation possible.</li></ul><figure><figcaption><strong>First-fifteen-minute response drill</strong><p>Times are coordination targets for the drill, not reported timings from a specific outage.</p></figcaption><table><thead><tr><th scope="col">Window</th><th scope="col">Coordinator</th><th scope="col">Technical lead</th></tr></thead><tbody><tr><th scope="row">0-3 minutes</th><td>State which flow is failing. Assign one person who can approve a change.</td><td>Find the last message that reached its expected business outcome.</td></tr><tr><th scope="row">3-7 minutes</th><td>Pause unrelated changes. Keep one shared timeline for the response.</td><td>Trace one message forward and mark the first checkpoint where evidence disappears.</td></tr><tr><th scope="row">7-12 minutes</th><td>Ask for one reversible mitigation and the signal that would stop it.</td><td>Apply the smallest safe action to the affected path.</td></tr><tr><th scope="row">12-15 minutes</th><td>Publish one short update. State the next action and the time of the next update.</td><td>Confirm that the business outcome recovered, not merely the pod.</td></tr></tbody></table></figure></section>
<section id="diagnosis"><h2>Separate process health from flow health</h2><p>A readiness probe answers whether the connector should receive work. A liveness probe answers whether restarting it could restore progress. Those are different decisions. If liveness restarts a connector that is merely slow, the restart can trigger a consumer rebalance. During a rebalance, Kafka redistributes partitions between consumer instances, so useful processing may pause and the backlog may grow.</p><p>I use Dynatrace to find when the error began and which downstream call changed. I use Conduktor to see whether the consumer offset is still moving. Then I connect both views with one correlation ID and confirm the final business state. Neither tool can prove completion on its own.</p><ul><li>A flat producer rate with falling completion suggests the break is downstream of emission.</li><li>Growing lag with stable consumer membership points toward processing capacity or a blocked dependency.</li><li>Frequent group rebalances after pod restarts make probe configuration and shutdown behavior suspects.</li><li>Stable offsets with missing business acknowledgements move the investigation into processing or persistence.</li></ul><p><strong>Sources: </strong><a href="#reference-kubernetes-probes">Kubernetes</a>, <a href="#reference-dynatrace-tracing">Dynatrace</a></p></section>
<section id="read-only-evidence"><h2>Collect read-only evidence before changing the flow</h2><p>The sequence below is a sanitized reconstruction, not a production transcript. It narrows the failed boundary without changing traffic or message positions.</p><p>Every command requires approved read access. Search logs only with an opaque correlation ID. Keep Kafka credentials inside a protected client configuration instead of placing them on the command line.</p><div><section><p><small>Step 1</small></p><h3>Check whether Kubernetes has been restarting the connector</h3><p>Start with process stability. This command shows the connector pods and their restart counts. It is a fast way to see whether the platform has been repeatedly replacing the process.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>kubectl -n &#x27;&lt;namespace&gt;&#x27; get pods -l &#x27;app=&lt;connector&gt;&#x27;</code></pre><p><strong>Sanitized output</strong></p><pre><code>NAME               READY   STATUS    RESTARTS
&lt;connector-pod&gt;    1/1     Running   &lt;count&gt;</code></pre></figure><p><strong>What this means: </strong>An increasing restart count explains interruptions at the process boundary. A running pod with no restarts does not prove that messages are moving, so the investigation continues.</p></section><section><p><small>Step 2</small></p><h3>Find the known message in recent connector logs</h3><p>Next, search for one approved correlation ID. The goal is to confirm whether this connector recorded the message and which topic it attempted to use.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>kubectl -n &#x27;&lt;namespace&gt;&#x27; logs &#x27;deployment/&lt;connector&gt;&#x27; --since=15m --prefix | grep -- &#x27;&lt;correlation-id&gt;&#x27;</code></pre><p><strong>Sanitized output</strong></p><pre><code>&lt;pod&gt; &lt;timestamp&gt; correlation_id=&lt;correlation-id&gt; published topic=&lt;topic&gt;</code></pre></figure><p><strong>What this means: </strong>A publish record moves the investigation toward the broker and consumer. No matching record keeps the investigation upstream, but only after confirming that the time window and log retention are sufficient.</p></section><section><p><small>Step 3</small></p><h3>Check whether the Kafka consumer group is advancing</h3><p>Finally, inspect the consumer-group position. CURRENT-OFFSET is the position saved by the consumer. LOG-END-OFFSET is the newest position in the partition. LAG is the difference between them.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>bin/kafka-consumer-groups.sh --bootstrap-server &#x27;&lt;broker&gt;&#x27; --command-config &#x27;&lt;protected-client.properties&gt;&#x27; --describe --group &#x27;&lt;consumer-group&gt;&#x27;</code></pre><p><strong>Sanitized output</strong></p><pre><code>TOPIC    PARTITION    CURRENT-OFFSET    LOG-END-OFFSET    LAG           CONSUMER-ID
&lt;topic&gt;  &lt;partition&gt;  &lt;committed&gt;       &lt;latest&gt;          &lt;difference&gt;  &lt;consumer-id&gt;</code></pre></figure><p><strong>What this means: </strong>Run this check again after a short observation window. A nonzero lag can be normal during a burst. If CURRENT-OFFSET does not advance, investigate the broker-to-consumer boundary. If it advances while the business result is missing, continue into processing or persistence.</p></section></div><p><strong>Sources: </strong><a href="#reference-kubernetes-logs">Kubernetes</a>, <a href="#reference-kafka-operations">Apache Kafka</a></p></section>
<section id="recovery"><h2>Recovery includes reconciliation and a tested way back</h2><p>A migration incident is not over when lag returns to zero. The consumer may have processed a message twice, or the producer may have skipped it before the broker ever saw it. I close the gap with a business reconciliation that finds missing and duplicate outcomes without exposing sensitive payloads.</p><p>Failback needs the same care as failover. First, confirm which environment is allowed to write. Next, prove that its checkpoint maps to the message position you expect. Keep mirroring active until the recovered path has handled normal and retry traffic for the agreed observation window. Removing it earlier can erase the safest route back.</p><figure><figcaption><strong>Exit conditions before closing the incident</strong><p>The exact thresholds belong to the service owner and its service-level objective; these dimensions should not be skipped.</p></figcaption><table><thead><tr><th scope="col">Dimension</th><th scope="col">Recovery question</th><th scope="col">Evidence</th></tr></thead><tbody><tr><th scope="row">Traffic</th><td>Is expected work entering the authoritative path?</td><td>Producer and ingress rates by flow</td></tr><tr><th scope="row">Progress</th><td>Is the oldest queued item getting newer rather than older?</td><td>Watch queue age while confirming that offsets move.</td></tr><tr><th scope="row">Correctness</th><td>Did work complete once and in the right state?</td><td>Domain reconciliation</td></tr><tr><th scope="row">Capacity</th><td>Can the recovered path absorb retries without falling behind again?</td><td>Drain a controlled backlog and watch saturation.</td></tr><tr><th scope="row">Reversal</th><td>Can we still return safely if the symptom recurs?</td><td>Run the route check and confirm which side owns writes.</td></tr></tbody></table></figure></section>
<section id="lessons"><h2>What this experience changed in my practice</h2><p>The migration reinforced that cross-environment reliability is mostly an ownership and evidence problem. Mirroring creates a return path, but it also creates more states for responders to understand. I now keep a one-page record for each important flow. The first line says which path is active and who owns it. The next line contains the query that proves the business outcome.</p><p>Private thresholds and incident results are intentionally omitted, so this article does not claim an improvement in mean time to recovery. The useful result is the operating model. A future lab can demonstrate how lag and redelivery behave during a safe reversal with synthetic data.</p><ul><li>Name the authoritative path for each stage of the migration.</li><li>Define recovery at the business boundary, not at the process boundary.</li><li>Use transport evidence to locate the break. Use business reconciliation to prove recovery.</li><li>Test failback while the original path is still healthy enough to teach you.</li></ul></section><section><h2>Key takeaways</h2><ul><li>A live cloud migration needs an evidence map for every critical message flow.</li><li>The first response objective is to name the affected flow and find the first boundary without evidence.</li><li>A healthy pod does not prove that a message completed. Verify the business outcome separately.</li><li>Mirroring is useful only when the team has tested the return route and knows which environment may write.</li><li>Honest anonymization is more credible than invented incident metrics.</li></ul></section><section><h2>References &amp; revisions</h2><ol><li id="reference-kafka-mirrormaker"><a href="https://kafka.apache.org/43/operations/">MirrorMaker and geo-replication documentation</a> — Apache Kafka</li><li id="reference-kubernetes-probes"><a href="https://kubernetes.io/docs/concepts/workloads/pods/probes/">Liveness, readiness, and startup probes</a> — Kubernetes</li><li id="reference-dynatrace-tracing"><a href="https://docs.dynatrace.com/docs/observe/application-observability/distributed-tracing">Distributed tracing concepts and analysis</a> — Dynatrace</li><li id="reference-azure-migrate"><a href="https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/migrate/plan-migration">Cloud Adoption Framework migration guidance</a> — Microsoft Azure</li><li id="reference-kubernetes-logs"><a href="https://kubernetes.io/docs/reference/kubectl/generated/kubectl_logs/">kubectl logs reference</a> — Kubernetes</li><li id="reference-kafka-operations"><a href="https://kafka.apache.org/43/operations/basic-kafka-operations/">Basic Kafka operations and consumer-group inspection</a> — Apache Kafka</li></ol><h3>Revision history</h3><p><time datetime="2026-08-25">2026-08-25</time> — Rebuilt the article around a step-by-step evidence path. Its read-only commands now use compact, copyable Automexia captures with separate sanitized output.</p></section>]]></content:encoded>
      </item>
      <item>
        <title>How I designed observability around a reactive payment journey</title>
        <link>https://amjedallaya.com/blog/observability-that-explains-production</link>
        <guid isPermaLink="true">https://amjedallaya.com/blog/observability-that-explains-production</guid>
        <pubDate>Fri, 24 Jul 2026 00:00:00 GMT</pubDate>
        <atom:updated>2026-08-25T00:00:00Z</atom:updated>
        <dc:creator>Amjed Allaya</dc:creator>
        <description>How I followed a payment from request to final state so dashboards explain user impact and alerts lead responders to the next useful check.</description>
        <category>Observability</category><category>Prometheus</category><category>Grafana</category><category>SLIs</category><category>RabbitMQ</category><category>Spring WebFlux</category><category>OpenTelemetry</category>
        <media:content url="https://amjedallaya.com/images/blog/observability-cover.webp" type="image/webp" medium="image" width="1672" height="941">
          <media:title>How I designed observability around a reactive payment journey</media:title>
          <media:description type="plain">An operator follows a failed payment from user impact to the responsible service.</media:description>
        </media:content>
        <content:encoded><![CDATA[<p><small>Published <time datetime="2026-07-24">July 24, 2026</time> · Updated <time datetime="2026-08-25">August 25, 2026</time></small></p><figure><img src="https://amjedallaya.com/images/blog/observability-cover.webp" alt="An operator follows a failed payment from user impact to the responsible service." width="1672" height="941"><figcaption>Telemetry becomes useful when it follows the same journey as the payment state.</figcaption></figure><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/observability-service-objectives.webp"/><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/observability-actionable-alerts.webp"/>
<div><p>A payment API can return a successful response because it accepted the work. The customer can still fail later if the queued message stops moving or the provider callback is rejected. That is why I do not treat HTTP availability as payment reliability.</p><p>I worked on Spring Boot WebFlux services that handed payment work to RabbitMQ and stored the final state in PostgreSQL. Redis reduced repeated reads, while Prometheus and Grafana exposed production behavior. The useful observability model followed the payment through those components instead of showing each component in isolation.</p></div>
<section id="journey"><h2>Start with the payment state machine</h2><p>The first dashboard should not begin with CPU or pod count. It should begin when the platform accepts the payment and end when the customer can see a final result. That result is a terminal state: a status such as completed or rejected where no more processing is expected. Between acceptance and that state, each change emits an event with a timestamp and outcome. A safe correlation value lets the operator follow one payment without exposing payment data.</p><figure><img src="https://amjedallaya.com/images/blog/observability-service-objectives.webp" alt="A payment journey is measured from acceptance to the result shown to the customer." width="1536" height="1024"/><figcaption>The service objective should describe the outcome the customer receives.</figcaption></figure><p>The decision was not to discard the HTTP indicator. It was to stop treating it as the main reliability result. A 202 response means the platform accepted work; it does not prove that the work completed. For this asynchronous path, the primary service-level indicator, or SLI, measures the proportion of valid attempts that reach an allowed terminal state inside the promised time.</p><figure><figcaption><strong>Payment-journey observability model</strong><p>Infrastructure signals remain diagnostic; the top-level objective follows the customer-visible state.</p></figcaption><table><thead><tr><th scope="col">Journey boundary</th><th scope="col">User-relevant question</th><th scope="col">Primary evidence</th></tr></thead><tbody><tr><th scope="row">API acceptance</th><td>Did the platform accept one valid request?</td><td>Compare accepted attempts with internal rejections.</td></tr><tr><th scope="row">Queue handoff</th><td>Did the command enter durable processing?</td><td>Confirm the publish, then watch the age of the oldest message.</td></tr><tr><th scope="row">Provider interaction</th><td>Did the external operation finish?</td><td>Record its final outcome and elapsed time.</td></tr><tr><th scope="row">Webhook processing</th><td>Was the callback authentic and handled once?</td><td>Record signature validation and the deduplication decision.</td></tr><tr><th scope="row">Terminal state</th><td>Can the customer see the correct result?</td><td>Read the domain state and measure time from acceptance to completion.</td></tr></tbody></table></figure></section>
<section id="signals"><h2>Give each telemetry signal one job</h2><p>Metrics show whether a group of payments is failing. Logs explain one decision, such as why a signature was rejected. Traces show where one payment spent its time as it crossed service boundaries. A deployment annotation adds the final question: what changed just before the failure began?</p><p>Correlation must survive the queue without leaking payment data. I use an opaque payment reference and keep card data or authentication material out of labels. A trace ID is valuable when investigating one request, but it has high cardinality: almost every request has a different value. Using it as a Prometheus label would create too many time series, so it belongs in logs and traces instead.</p><figure><figcaption><strong>Automexia code terminal — Sanitized payment event shape</strong> <small>JSON</small></figcaption><pre><code>{
  &quot;timestamp&quot;: &quot;2026-08-25T10:42:18Z&quot;,
  &quot;event.name&quot;: &quot;payment.webhook.processed&quot;,
  &quot;service.name&quot;: &quot;payment-webhook-handler&quot;,
  &quot;service.version&quot;: &quot;&lt;release-id&gt;&quot;,
  &quot;deployment.environment.name&quot;: &quot;production&quot;,
  &quot;trace_id&quot;: &quot;&lt;opaque-trace-id&gt;&quot;,
  &quot;payment_reference&quot;: &quot;&lt;opaque-domain-reference&gt;&quot;,
  &quot;provider&quot;: &quot;&lt;bounded-provider-name&gt;&quot;,
  &quot;outcome&quot;: &quot;accepted&quot;,
  &quot;duration_ms&quot;: 84
}</code></pre><p>The shape is sanitized and aligned with OpenTelemetry naming where an applicable stable convention exists. Sensitive payment values are deliberately absent.</p></figure><p><strong>Sources: </strong><a href="#reference-otel-semconv">OpenTelemetry</a></p></section>
<section id="slis"><h2>Define SLIs that cannot be fooled by HTTP success</h2><p>I use one SLI for acceptance and another for completion. The acceptance SLI tells me whether the synchronous API can take valid work. The completion SLI tells me whether that work reaches an allowed final state before the promised deadline. Keeping them separate shows whether the failure happened before or after the queue handoff.</p><p>Correctness is checked at the final domain state. Callback signature failures remain a security signal, but rejecting an invalid callback is expected behavior rather than an availability failure.</p><p>The query below shows the shape of a completion ratio. It assumes that every eligible attempt produces one terminal observation. A reconciler marks work as timed out after its deadline, so stuck work cannot disappear from the calculation. Before adopting the query, a team must write down what makes an attempt eligible and which final states count as success.</p><ul><li>Acceptance: valid payment commands accepted without an internal error.</li><li>Completion: accepted commands reaching a permitted terminal state within the objective.</li><li>Latency: completion duration measured end to end, not only controller response time.</li><li>Correctness: confirm that the stored result matches the provider result. Then verify that the customer sees that same state.</li></ul><figure><figcaption><strong>Automexia code terminal — Sanitized completion SLI</strong> <small>PromQL</small></figcaption><pre><code>sum(rate(payment_outcomes_total{
  eligible=&quot;true&quot;,
  outcome=&quot;completed&quot;,
  within_slo=&quot;true&quot;
}[30m]))
/
sum(rate(payment_outcomes_total{eligible=&quot;true&quot;}[30m]))</code></pre><p>Before using this query, enforce one terminal observation per attempt. Then define how late completion and zero traffic affect the result.</p></figure></section>
<section id="investigation"><h2>Make the dashboard an investigation path</h2><p>The top row answers whether customers are receiving a result on time. It shows how many accepted payments finish successfully. Beside that result, the age of the oldest unfinished payment reveals whether delayed work is still progressing. The next row shows where that work is waiting. Infrastructure charts appear only after the dashboard has located the affected stage.</p><p>That order matters during an incident. A responder can see that acceptance is stable while completion falls and queue age rises. The queue-age chart should open the queue view for that service. A webhook failure should open filtered traces for that handler. Each link continues the investigation instead of sending the reader to a generic platform dashboard.</p><figure><figcaption><strong>Symptom-to-evidence investigation path</strong><p>Each row narrows the next question and avoids searching all telemetry at once.</p></figcaption><table><thead><tr><th scope="col">Observed symptom</th><th scope="col">Next comparison</th><th scope="col">Likely investigation branch</th></tr></thead><tbody><tr><th scope="row">Acceptance falls</th><td>Compare errors before and after the latest release.</td><td>Start at the API path and its immediate dependency.</td></tr><tr><th scope="row">Acceptance is stable while queue age rises</th><td>Compare the publish rate with the consumer completion rate.</td><td>Trace one consumer job to find blocking work or a slow dependency.</td></tr><tr><th scope="row">Webhook rejection rises</th><td>Group rejection by signature result.</td><td>Check whether configuration changed before inspecting provider payload shape.</td></tr><tr><th scope="row">Completion slows while service resources remain stable</th><td>Compare provider time with database time.</td><td>Investigate the slower external boundary.</td></tr></tbody></table></figure></section>
<section id="alerts"><h2>Page on threatened outcomes, then attach causes</h2><p>A useful page begins with one sentence about user impact. The alert goes to the team that can change the failing service. Its link opens the dashboard at the threatened objective, not a generic infrastructure page. The runbook starts with one safe check that can narrow the problem. Queue depth alone cannot do this because the same depth may be normal during a burst. The age of the oldest item reveals whether work is still progressing.</p><figure><img src="https://amjedallaya.com/images/blog/observability-actionable-alerts.webp" alt="A failed user journey triggers one alert connected to a focused dashboard and a safe runbook action." width="1536" height="1024"/><figcaption>A page should connect a threatened objective to an action a responder can take now.</figcaption></figure><p>I prefer multi-window alerting. One short window detects a severe failure quickly, while a longer window catches a smaller failure that continues. The burn rate describes how quickly the service is consuming its allowed error budget. Its threshold and evaluation windows must come from the service objective and must be tested with historical or synthetic events. Related cause alerts should be grouped or inhibited so one provider failure does not wake every downstream team independently.</p><ul><li>Page: a payment objective is burning fast enough to require immediate human action.</li><li>Ticket: an operational weakness needs work but is not harming users now.</li><li>Dashboard annotation: mark a production change so later symptoms have a clear comparison point.</li><li>Security signal: signature failures require separate routing and should not expose payload data.</li></ul><p><strong>Sources: </strong><a href="#reference-prometheus-alerts">Prometheus</a></p></section>
<section id="test-alerts"><h2>Test an alert before it can page anyone</h2><p>A syntax check only proves that Prometheus can parse the rule. It does not prove that the alert waits for the intended duration or clears when the payment journey recovers. A small unit-test file should live beside the rule. The test feeds synthetic metric values into the expression and states exactly when the alert should be silent and when it should fire.</p><p>The useful minimum is three cases. Normal traffic must remain silent. A brief spike shorter than the configured duration must also remain silent. A sustained completion failure must fire and identify the owning team. It must also include the runbook link. This makes the alert behavior reviewable before it can interrupt a person.</p><div><section><p><small>Step 1</small></p><h3>Check that Prometheus can parse the alert rule</h3><p>Run the static check before loading the rule anywhere. It catches malformed YAML and invalid rule definitions without contacting the production Prometheus server.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>promtool check rules payment-alerts.yml</code></pre><p><strong>Sanitized output</strong></p><pre><code>Checking payment-alerts.yml
  SUCCESS: &lt;rule-count&gt; rules found</code></pre></figure><p><strong>What this means: </strong>Success means the file is structurally valid. It says nothing about whether the alert fires at the correct time, so a behavior test is still required.</p></section><section><p><small>Step 2</small></p><h3>Replay the alert against synthetic metric values</h3><p>Now evaluate the rule with the unit-test file. The test file should describe normal traffic, a short spike, and a sustained failure without using customer data.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>promtool test rules payment-alerts.test.yml</code></pre><p><strong>Sanitized output</strong></p><pre><code>Unit Testing: payment-alerts.test.yml
  SUCCESS</code></pre></figure><p><strong>What this means: </strong>Success means the declared synthetic cases produced the expected alert states. It does not prove that production labels are correct or that a notification reaches the on-call engineer. Those require a controlled end-to-end test.</p></section></div><p><strong>Sources: </strong><a href="#reference-prometheus-rule-tests">Prometheus</a></p></section>
<section id="limits"><h2>State what the evidence proves—and what it does not</h2><p>I can substantiate two production facts. The reactive services were designed for more than 500 transactions per minute. Prometheus and Grafana were introduced so operators could see service behavior and diagnose incidents. I do not have a publishable before-and-after measurement for mean time to detect or an alert-volume baseline, so I will not invent one.</p><p>The next step is to version the telemetry contract so a renamed field cannot silently break a dashboard. I would also trigger each critical alert during a controlled failure. That test should prove that the alert reaches the correct team and opens useful evidence. Finally, I would track how many payments lose correlation. This shows whether an investigation can still follow one payment from start to finish.</p><p>Observability also has a cost. Dividing telemetry spend by successful payments makes that cost visible as traffic grows.</p><p>A histogram places latency observations into configured ranges called buckets. Those boundaries need to sit around the latency objective the service actually promises. Averaging percentiles from individual instances does not produce a valid percentile for the fleet.</p><ul><li>Keep a bounded label policy and review new dimensions before deployment.</li><li>Preserve errors and slow traces intentionally when sampling normal traffic.</li><li>Monitor the monitoring path. Alert when telemetry arrives late or disappears.</li><li>Record which OpenTelemetry naming version the service follows because field names evolve.</li></ul><p><strong>Sources: </strong><a href="#reference-prometheus-histograms">Prometheus</a>, <a href="#reference-grafana-dashboards">Grafana Labs</a></p></section><section><h2>Key takeaways</h2><ul><li>Model observability around the payment state machine, not the deployment diagram.</li><li>Separate request acceptance from asynchronous completion and correctness.</li><li>Keep high-cardinality and sensitive values out of metric labels.</li><li>Make dashboards and alerts lead from user impact to one bounded investigation path.</li><li>Publish measured results and sanitized examples as different kinds of evidence.</li></ul></section><section><h2>References &amp; revisions</h2><ol><li id="reference-otel-semconv"><a href="https://opentelemetry.io/docs/specs/semconv/">OpenTelemetry semantic conventions</a> — OpenTelemetry</li><li id="reference-prometheus-histograms"><a href="https://prometheus.io/docs/practices/histograms/">Histograms and summaries</a> — Prometheus</li><li id="reference-prometheus-alerts"><a href="https://prometheus.io/docs/prometheus/latest/configuration/alerting_rules/">Alerting rules</a> — Prometheus</li><li id="reference-grafana-dashboards"><a href="https://grafana.com/docs/grafana/latest/visualizations/dashboards/">Dashboard documentation</a> — Grafana Labs</li><li id="reference-prometheus-rule-tests"><a href="https://prometheus.io/docs/prometheus/latest/configuration/unit_testing_rules/">Unit testing for Prometheus rules</a> — Prometheus</li></ol><h3>Revision history</h3><p><time datetime="2026-08-25">2026-08-25</time> — Rebuilt the article around a guided payment investigation. Its Prometheus checks now use compact, copyable Automexia captures with colored validation output.</p></section>]]></content:encoded>
      </item>
      <item>
        <title>From two hours to ten minutes: automating deployment safely</title>
        <link>https://amjedallaya.com/blog/safe-automation-reduces-human-error</link>
        <guid isPermaLink="true">https://amjedallaya.com/blog/safe-automation-reduces-human-error</guid>
        <pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate>
        <atom:updated>2026-08-25T00:00:00Z</atom:updated>
        <dc:creator>Amjed Allaya</dc:creator>
        <description>How I reduced a two-hour manual deployment to about ten minutes by making every release stage explain why it could continue or had to stop.</description>
        <category>Automation</category><category>Azure DevOps</category><category>CI/CD</category><category>Kubernetes</category><category>Helm</category><category>Deployment safety</category>
        <media:content url="https://amjedallaya.com/images/blog/automation-cover.webp" type="image/webp" medium="image" width="1672" height="941">
          <media:title>From two hours to ten minutes: automating deployment safely</media:title>
          <media:description type="plain">An operator turns a manual deployment checklist into a controlled workflow.</media:description>
        </media:content>
        <content:encoded><![CDATA[<p><small>Published <time datetime="2026-06-30">June 30, 2026</time> · Updated <time datetime="2026-08-25">August 25, 2026</time></small></p><figure><img src="https://amjedallaya.com/images/blog/automation-cover.webp" alt="An operator turns a manual deployment checklist into a controlled workflow." width="1672" height="941"><figcaption>The useful automation result is a faster decision path with stronger evidence, not merely fewer clicks.</figcaption></figure><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/automation-contract.webp"/><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/automation-failure-paths.webp"/>
<div><p>The deployment problem was not that engineers were slow. They had to remember which artifact had passed testing. They also had to reconstruct the correct environment settings before deciding whether the rollout was healthy. None of those decisions lived in one reliable record. Repeating the process manually made a routine release take roughly two hours.</p><p>I helped move that workflow into Azure DevOps and reduced the path to roughly ten minutes. The pipeline built one identified artifact and carried it through each environment. Every stage answered a clear question before the release could continue. The speed came from removing repeated reasoning, not from skipping controls.</p></div>
<section id="before"><h2>Map decisions before automating commands</h2><p>A manual runbook mixes repeatable commands with human judgment. Compiling the application and running the same test suite should produce the same result every time, so the pipeline should own that work. Deciding whether a database change remains backward compatible requires context and an accountable human.</p><p>For each stage, I first wrote what it accepts. I then wrote the result it must produce and the condition that should stop it. Finally, I named the person who can approve an exception. Automating an undocumented sequence would have made it faster without making it safer.</p><p>I rejected the idea of hiding the whole runbook inside one large script. That design would make a partial failure difficult to locate and would leave a reviewer unsure about what had already changed. Named stages preserve the decision points. They also let the pipeline stop before production access is granted.</p><figure><figcaption><strong>Manual responsibility versus pipeline responsibility</strong><p>The boundary keeps high-value judgment visible while removing repeated mechanics.</p></figcaption><table><thead><tr><th scope="col">Concern</th><th scope="col">Pipeline owns</th><th scope="col">Human owns</th></tr></thead><tbody><tr><th scope="row">Artifact</th><td>Build once and attach an immutable identity to the output.</td><td>Approve that exact release candidate.</td></tr><tr><th scope="row">Quality</th><td>Run each test layer once at the stage where it adds evidence.</td><td>Approve only an exception that is written and reviewable.</td></tr><tr><th scope="row">Target</th><td>Bind deployment to a named protected environment.</td><td>Authorize production access and choose the release time.</td></tr><tr><th scope="row">Rollout</th><td>Apply the Helm change and report Kubernetes progress.</td><td>Choose rollback when the result is unclear.</td></tr><tr><th scope="row">Verification</th><td>Run a health check and one critical smoke test. A smoke test is a small check that proves the most important path can start and finish.</td><td>Decide whether the business needs wider validation.</td></tr></tbody></table></figure></section>
<section id="contract"><h2>Make the delivery contract visible in stages</h2><p>Stages are useful because they create named evidence boundaries. A production deployment should never rebuild an artifact that passed tests earlier; it should promote the identified artifact. Environment-specific configuration is injected at deployment time from controlled resources, not copied into the repository or typed into a terminal.</p><figure><img src="https://amjedallaya.com/images/blog/automation-contract.webp" alt="An operator reviews a deployment as it passes through controlled stages." width="1672" height="941"/><figcaption>A safe pipeline shows what entered each stage and why it was allowed to continue.</figcaption></figure><p>The skeleton below communicates the dependency graph. It is intentionally incomplete: production approvals and checks belong to protected Azure DevOps environment resources, where editing pipeline YAML alone cannot silently remove them.</p><figure><figcaption><strong>Automexia code terminal — Simplified multi-stage delivery shape</strong> <small>YAML</small></figcaption><pre><code>stages:
  - stage: Validate
    jobs: [unit_tests, integration_tests, security_checks]

  - stage: Package
    dependsOn: Validate
    jobs: [build_once, publish_artifact, render_helm]

  - stage: Deploy_Staging
    dependsOn: Package
    jobs: [deploy_exact_artifact, smoke_test]

  - stage: Deploy_Production
    dependsOn: Deploy_Staging
    environment: production
    jobs: [deploy_bounded_change, verify_rollout]

  - stage: Verify_Outcome
    dependsOn: Deploy_Production
    jobs: [journey_check, publish_release_evidence]</code></pre><p>This is a sanitized design sketch, not executable production YAML. The operational details are deliberately omitted.</p></figure><p><strong>Sources: </strong><a href="#reference-azure-pipelines">Microsoft Learn</a></p></section>
<section id="gates"><h2>Use gates that answer different failure questions</h2><p>More tests do not automatically create a safer release. Each gate should rule out a different class of failure. Unit tests protect local decision logic. WireMock and Testcontainers exercise integration behavior with realistic dependencies. Contract tests protect service boundaries. Playwright protects a small number of critical journeys. Re-running identical tests at every stage only increases time without adding evidence.</p><p>A gate must also fail clearly. The release record should name the failed check and the exact release that was running. It must also say whether production changed before the failure. A red stage with thousands of ungrouped log lines recreates the same cognitive load as the manual process.</p><ul><li>Fast gate: rejects code that cannot compile or pass its unit tests.</li><li>Dependency gate: starts isolated dependencies and proves that the service can communicate with them.</li><li>Contract gate: proves that a new build still speaks the interface expected by existing callers.</li><li>Journey gate: runs a small synthetic payment through the behavior users depend on.</li><li>Deployment gate: shows the exact image and configuration change before promotion.</li></ul></section>
<section id="permissions"><h2>Keep production authority outside the pipeline file</h2><p>A contributor who can edit YAML should not automatically gain the authority to deploy to production. I attach approval rules to a protected Azure DevOps environment. The environment is a separately controlled deployment target, so its policy does not live in the pipeline file. A pull request can change the workflow, but it cannot silently remove the policy that protects production.</p><p>The service connection is scoped to the environment it must change. Secrets stay in protected storage and appear in the pipeline only as references. An approval is useful when the reviewer can see the exact artifact and configuration difference. Without that evidence, the approval only adds waiting time.</p><figure><figcaption><strong>Controls that should not rely on convention</strong><p>A documented rule is weaker than a technical boundary that records and rejects violations.</p></figcaption><table><thead><tr><th scope="col">Risk</th><th scope="col">Technical control</th><th scope="col">Evidence retained</th></tr></thead><tbody><tr><th scope="row">Wrong environment</th><td>Use a protected environment with a scoped service connection.</td><td>Record who requested and approved the target.</td></tr><tr><th scope="row">Wrong artifact</th><td>Promote an immutable digest between stages.</td><td>Link the digest to its source commit.</td></tr><tr><th scope="row">Concurrent releases</th><td>Allow only one production deployment at a time.</td><td>Show which run waited or was replaced.</td></tr><tr><th scope="row">Secret exposure</th><td>Read secrets from a protected store and mask logs.</td><td>Record the secret reference, never its value.</td></tr><tr><th scope="row">Policy bypass</th><td>Keep mandatory checks outside editable YAML.</td><td>Retain the approval history.</td></tr></tbody></table></figure><p><strong>Sources: </strong><a href="#reference-azure-approvals">Microsoft Learn</a>, <a href="#reference-azure-security">Microsoft Learn</a></p></section>
<section id="verify-artifact"><h2>Verify the image before checking the rollout</h2><p>An image tag such as release-latest can move to different content. An image digest is derived from the image content and identifies one exact image. The pipeline should therefore carry the approved digest into production and compare it with the image declared on the Deployment.</p><p>Both commands below are read-only and require namespace-scoped access. The image comparison happens before rollout validation because a healthy rollout of the wrong artifact is still a failed release.</p><div><section><p><small>Step 1</small></p><h3>Read the exact image declared on the Deployment</h3><p>Ask Kubernetes for the image assigned to the application container. In a multi-container pod, the container name must be explicit so a sidecar image is not compared by mistake.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>kubectl -n &#x27;&lt;namespace&gt;&#x27; get &#x27;deployment/&lt;service&gt;&#x27; -o jsonpath=&#x27;{.spec.template.spec.containers[?(@.name==&quot;&lt;container&gt;&quot;)].image}&#x27;</code></pre><p><strong>Sanitized output</strong></p><pre><code>&lt;registry&gt;/&lt;service&gt;@sha256:&lt;deployed-digest&gt;</code></pre></figure><p><strong>What this means: </strong>Compare the complete output with the digest approved earlier in the pipeline. Any mismatch stops the release. A tag-only value is also a failure because it does not identify immutable content.</p></section><section><p><small>Step 2</small></p><h3>Wait for Kubernetes to finish the rollout</h3><p>Only after the digest matches should the pipeline wait for the Deployment rollout. The timeout must reflect how long this service normally needs to become ready.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>kubectl -n &#x27;&lt;namespace&gt;&#x27; rollout status &#x27;deployment/&lt;service&gt;&#x27; --timeout=&#x27;&lt;agreed-timeout&gt;&#x27;</code></pre><p><strong>Sanitized output</strong></p><pre><code>deployment &quot;&lt;service&gt;&quot; successfully rolled out</code></pre></figure><p><strong>What this means: </strong>This proves that Kubernetes completed its rollout checks. It does not prove that a customer journey works. The next stage still needs the critical smoke test, recorded against the same digest.</p></section></div><p><strong>Sources: </strong><a href="#reference-kubernetes-images">Kubernetes</a>, <a href="#reference-kubernetes-rollout">Kubernetes</a></p></section>
<section id="failure-path"><h2>Design rollback and partial failure before promotion</h2><p>Retries are safe for reads and for steps that are genuinely idempotent. An idempotent operation reaches the same state when it is repeated, so a lost response does not turn the retry into a second change. Retries are dangerous when a write may have succeeded but its response was lost. Before repeating a deployment task, the pipeline reads the current external state and checks whether the declared artifact is already active.</p><figure><img src="https://amjedallaya.com/images/blog/automation-failure-paths.webp" alt="A failed automated task branches into retry, rollback, or human intervention." width="1536" height="1024"/><figcaption>A pipeline is trustworthy when it explains the safest next action after partial failure.</figcaption></figure><p>Helm keeps release revisions, but rollback is not a universal undo. A database migration may have changed data that the old application cannot read. I therefore expand the schema before using new fields and postpone removal until the old release is outside the rollback window. The same compatibility question applies to messages and external side effects.</p><ul><li>Retry: a read-only step is usually safe to repeat. A write must be idempotent or protected by an idempotency key before the pipeline retries it.</li><li>Rollback: when the previous application and configuration remain compatible with current state.</li><li>Compensate: when an external side effect cannot be erased but an explicit counter-action exists.</li><li>Stop for a human: when the pipeline cannot prove what state production is in.</li></ul><p><strong>Sources: </strong><a href="#reference-helm-rollback">Helm</a>, <a href="#reference-kubernetes-deployment">Kubernetes</a></p></section>
<section id="result"><h2>Measure the result without overclaiming it</h2><p>The result I can substantiate is deployment duration: approximately two hours manually and approximately ten minutes through the automated path, a reduction of roughly 92 percent. That improvement matters because feedback arrives in the same working session and the release no longer depends on reconstructing a long sequence from memory.</p><p>I do not have a publishable baseline for change-failure rate, so I will not claim a reliability improvement by percentage. The next version should measure how long each release waits and how long it executes. It should also record when a person intervenes or a post-deployment incident follows.</p><figure><figcaption><strong>Measured deployment duration</strong><p>Portfolio result from the payment-platform delivery workflow. Values are approximate end-to-end durations.</p></figcaption><table><tbody><tr><th scope="row">Manual workflow</th><td>~120 minutes</td></tr><tr><th scope="row">Automated workflow</th><td>~10 minutes</td></tr></tbody></table></figure><blockquote>Automation earned trust because it made the release state easier to inspect. The speed was the consequence.</blockquote></section><section><h2>Key takeaways</h2><ul><li>Automate deterministic mechanics while keeping accountable judgment explicit.</li><li>Build once and promote the same immutable artifact through every environment.</li><li>Protect production with resource-owned permissions and checks outside editable YAML.</li><li>Retry only after proving what the previous attempt changed. Stop when production state is unclear.</li><li>Report the measured duration improvement without inventing reliability percentages.</li></ul></section><section><h2>References &amp; revisions</h2><ol><li id="reference-azure-pipelines"><a href="https://learn.microsoft.com/en-us/azure/devops/pipelines/get-started/key-pipelines-concepts?view=azure-devops">Key Azure Pipelines concepts</a> — Microsoft Learn</li><li id="reference-azure-approvals"><a href="https://learn.microsoft.com/en-us/azure/devops/pipelines/process/approvals?view=azure-devops">Approvals and checks</a> — Microsoft Learn</li><li id="reference-azure-security"><a href="https://learn.microsoft.com/en-us/azure/devops/pipelines/policies/permissions?view=azure-devops">Manage security in Azure Pipelines</a> — Microsoft Learn</li><li id="reference-helm-rollback"><a href="https://helm.sh/docs/intro/using_helm/#helm-upgrade-and-helm-rollback-upgrading-a-release-and-recovering-on-failure">Helm upgrade and rollback</a> — Helm</li><li id="reference-kubernetes-deployment"><a href="https://kubernetes.io/docs/concepts/workloads/controllers/deployment/">Kubernetes Deployments</a> — Kubernetes</li><li id="reference-kubernetes-images"><a href="https://kubernetes.io/docs/concepts/containers/images/">Container images and immutable digests</a> — Kubernetes</li><li id="reference-kubernetes-rollout"><a href="https://kubernetes.io/docs/reference/kubectl/generated/kubectl_rollout/kubectl_rollout_status/">kubectl rollout status reference</a> — Kubernetes</li></ol><h3>Revision history</h3><p><time datetime="2026-08-25">2026-08-25</time> — Rebuilt the article as a measured deployment case study. Image and rollout checks now use compact, copyable Automexia captures with separate sanitized output.</p></section>]]></content:encoded>
      </item>
      <item>
        <title>What two production workloads taught me about performance and capacity</title>
        <link>https://amjedallaya.com/blog/performance-resource-and-cloud-cost</link>
        <guid isPermaLink="true">https://amjedallaya.com/blog/performance-resource-and-cloud-cost</guid>
        <pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate>
        <atom:updated>2026-08-25T00:00:00Z</atom:updated>
        <dc:creator>Amjed Allaya</dc:creator>
        <description>How I measured two different workloads and avoided performance or cost claims that the data could not support.</description>
        <category>Performance &amp; capacity</category><category>Spring WebFlux</category><category>RabbitMQ</category><category>Redis</category><category>Kubernetes</category><category>Java performance</category>
        <media:content url="https://amjedallaya.com/images/blog/performance-cover.webp" type="image/webp" medium="image" width="1672" height="941">
          <media:title>What two production workloads taught me about performance and capacity</media:title>
          <media:description type="plain">An operator locates a database constraint before increasing worker capacity.</media:description>
        </media:content>
        <content:encoded><![CDATA[<p><small>Published <time datetime="2026-05-27">May 27, 2026</time> · Updated <time datetime="2026-08-25">August 25, 2026</time></small></p><figure><img src="https://amjedallaya.com/images/blog/performance-cover.webp" alt="An operator locates a database constraint before increasing worker capacity." width="1672" height="941"><figcaption>The useful optimization starts with the constrained workload, not the loudest resource graph.</figcaption></figure><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/performance-remove-work.webp"/><link rel="preload" as="image" href="https://amjedallaya.com/images/blog/performance-capacity-headroom.webp"/>
<div><p>“Make it faster” described two different problems in my work. A reactive payment path handled bursts and waited for external callbacks before showing a final result. A document engine used Quarkus and Apache POI to generate more than 7,000 documents per day.</p><p>For payments, speed meant reaching the correct final state before the customer noticed a delay. For document generation, speed meant draining the backlog without exhausting memory. The same CPU graph cannot explain both outcomes, so I start by describing the work before interpreting the resources.</p></div>
<section id="workload-shape"><h2>Describe the workload before interpreting utilization</h2><p>A performance baseline starts by defining one unit of completed work. For payments, the unit cannot stop at the HTTP response because the provider callback arrives later. For documents, the unit must include a correct file and an advanced workflow state.</p><p>I then measure how the work arrives. A steady rate behaves differently from a short burst, even when the daily total is the same. Finally, I record how many units run together and how payload size changes their service time. Those details make later CPU and memory graphs interpretable.</p><p>This prevents a common mistake: comparing average CPU between services and calling the lower number waste. A waiting service can have low CPU and poor latency; a compute-heavy batch worker can have high CPU and healthy throughput.</p><p>The decision was to stop asking which service used more resources and instead ask whether each workload completed safely. That changed the experiment. The payment test follows end-to-end completion during a burst, while the document test measures backlog progress and memory retained by each active job.</p><figure><figcaption><strong>Two workload shapes, two performance contracts</strong><p>The figures are public portfolio facts; the diagnostic dimensions are the operating model used for each workload class.</p></figcaption><table><thead><tr><th scope="col">Dimension</th><th scope="col">Reactive payment path</th><th scope="col">Document-generation path</th></tr></thead><tbody><tr><th scope="row">Published scale</th><td>&gt;500 transactions/minute design target</td><td>&gt;7,000 generated documents/day capability</td></tr><tr><th scope="row">Primary outcome</th><td>Correct terminal payment state inside a latency objective</td><td>Correct document produced and workflow advanced</td></tr><tr><th scope="row">Where time may go</th><td>The service may wait for a provider response or queued work.</td><td>The worker may wait for source data or storage.</td></tr><tr><th scope="row">Where pressure appears</th><td>Too many concurrent requests can exhaust a connection pool.</td><td>Large files and too many workers can retain excessive heap.</td></tr><tr><th scope="row">Capacity signal</th><td>Track unfinished work and the age of its oldest item.</td><td>Track backlog age and memory used by one active job.</td></tr></tbody></table></figure></section>
<section id="measurement"><h2>Find out where elapsed time is spent</h2><p>Elapsed time is a budget. Some of it is active application work. The rest may be spent waiting for a dependency or sitting in a queue before a worker starts. A trace explains an online request, while queue timestamps explain delay that happens before the process receives the job. A CPU profiler cannot see that queue wait.</p><p>The baseline template keeps the experiment honest. I warm the system first, then measure normal traffic and a representative peak. If deployments or scaling events are part of daily operation, the window includes one of them. This prevents a clean laboratory interval from hiding the behavior operators actually face.</p><figure><figcaption><strong>Automexia code terminal — Performance experiment record</strong> <small>YAML</small></figcaption><pre><code>workload: &lt;payment-or-document-flow&gt;
window: &lt;representative-period&gt;
input:
  arrival_shape: &lt;steady-burst-or-batch&gt;
  concurrency: &lt;measured-value&gt;
  payload_distribution: &lt;p50-p95-max&gt;
observe:
  - completion_rate
  - p50_p95_p99_or_job_duration
  - queue_age_and_depth
  - dependency_wait
  - cpu_memory_gc
  - retries_and_redelivery
change: &lt;one-controlled-variable&gt;
rollback_when: &lt;pre-agreed-condition&gt;</code></pre><p>The record is a reusable template. It deliberately contains no invented measurements.</p></figure></section>
<section id="reactive"><h2>Treat reactive code as a concurrency model, not a speed claim</h2><p>Spring WebFlux provides non-blocking I/O, but a reactive controller cannot make a blocking dependency disappear. A traditional JDBC database call or a synchronous library supplied by a provider can still occupy the event loop. When that happens, a small amount of blocking work can increase tail latency for many unrelated requests.</p><p>I inspect thread use at every boundary. Blocking work is either removed or moved to a bounded scheduler with a clear concurrency limit. During the load test, I first confirm that the event loop keeps making progress. Then I watch the scarce dependency pool and the age of unfinished work. I also measure p99 latency, the duration below which 99 percent of completed requests fall. Controller latency alone is not enough because it ends before the payment journey does.</p><ul><li>Propagate deadlines so abandoned requests do not continue consuming downstream capacity.</li><li>Bound concurrency before a scarce provider or database connection pool becomes the queue.</li><li>Measure p99 and timeout rate; averages hide the customers waiting longest.</li><li>Warm the system with realistic payloads. Verify that the load generator can produce the intended rate.</li></ul><p><strong>Sources: </strong><a href="#reference-spring-webflux">Spring Framework</a></p></section>
<section id="messaging"><h2>Tune queue consumers for safe recovery</h2><p>RabbitMQ prefetch controls how much unacknowledged work one consumer can hold. Raising it may improve throughput, but every extra message also consumes memory and may be replayed after a failure. The useful limit is therefore the amount of in-flight work the process and its downstream dependency can recover safely.</p><p>I create a controlled backlog and raise prefetch in small steps. After each step, I check whether the oldest message becomes younger and whether memory remains comfortably below the container limit. I then restart a consumer and observe how much work is redelivered. A higher throughput number is rejected if recovery becomes unsafe.</p><figure><figcaption><strong>Consumer signals and what they imply</strong><p>Interpret combinations rather than reacting to one queue graph.</p></figcaption><table><thead><tr><th scope="col">Observation</th><th scope="col">Likely constraint</th><th scope="col">Next experiment</th></tr></thead><tbody><tr><th scope="row">Backlog grows while downstream calls slow</th><td>The dependency is the limiting boundary.</td><td>Reduce concurrency and confirm that dependency latency recovers.</td></tr><tr><th scope="row">Backlog grows while CPU stays low</th><td>The consumer may be blocked or under-delivering work.</td><td>Trace one job before changing prefetch.</td></tr><tr><th scope="row">Throughput improves but memory jumps</th><td>The consumer is holding too much unfinished work.</td><td>Lower prefetch and repeat the failure test.</td></tr><tr><th scope="row">Restarting creates heavy redelivery</th><td>Acknowledgement may happen at the wrong boundary.</td><td>Acknowledge only after the durable outcome, then test replay again.</td></tr></tbody></table></figure><p><strong>Sources: </strong><a href="#reference-rabbitmq-confirms">RabbitMQ</a></p></section>
<section id="cache"><h2>Cache stable reads, then design invalidation and failure</h2><p>On the payment platform, Redis cached tokens and application configuration so the services repeated fewer PostgreSQL reads. I did not give both data types one generic time to live, or TTL. The TTL defines how long a cached value may remain before it expires. A token has a security boundary, while configuration has a different tolerance for staleness.</p><figure><img src="https://amjedallaya.com/images/blog/performance-remove-work.webp" alt="An operator removes repeated work before adding more capacity." width="1536" height="1024"/><figcaption>Removing repeated work is often safer than buying more capacity.</figcaption></figure><p>For each cached value, I ask how stale it may become. I name the event that invalidates it and decide what the service should do when Redis is unavailable. For security-sensitive configuration, a slower authoritative read may be safer than serving stale data.</p><p>The cache must also avoid a stampede after expiration. A stampede happens when many requests notice the same missing value and all reload it from the primary store. I coalesce those requests so one load supplies the others, then bound the memory used by cached entries. Hit rate is useful only when it is read beside primary-store load and correctness failures.</p><ul><li>Classify data by acceptable staleness and failure consequence.</li><li>Coalesce concurrent requests for the same hot key after expiration.</li><li>Track hit rate together with primary-store load and correctness incidents.</li><li>Flush or bypass local cache state when invalidation connectivity is uncertain.</li></ul><p><strong>Sources: </strong><a href="#reference-redis-cache">Redis</a></p></section>
<section id="capacity"><h2>Size Kubernetes memory from process behavior, not heap alone</h2><p>Kubernetes uses the memory request when scheduling a pod and enforces the memory limit while it runs. A Java process consumes more than heap. Thread stacks and direct buffers, for example, live outside the heap but still count toward the container limit. Setting the limit equal to the heap target removes the headroom those components need.</p><figure><img src="https://amjedallaya.com/images/blog/performance-capacity-headroom.webp" alt="An operator keeps memory headroom above the workload used by a Java service." width="1536" height="1024"/><figcaption>The container needs room for Java heap and the native memory used around it.</figcaption></figure><p>The decision I changed was to stop using the Java heap target as the container size. Heap describes memory used for Java objects, not the whole process. I start with resident memory, which is the process memory currently held in physical memory from the operating system perspective. I then compare it with heap after garbage collection and investigate the remaining non-heap use.</p><p>On the payment platform, tuning the application and container settings reduced OOMKill incidents. An OOMKill occurs when Kubernetes terminates a container after it exceeds the enforced memory boundary. I do not have a publishable incident count, so the result remains directional rather than a percentage improvement.</p><p>I change the JVM setting and Kubernetes limit separately when possible. After each change, I repeat the same workload and watch garbage-collection pauses as well as resident memory. This makes a regression easier to attribute and keeps the previous resource profile available for rollback.</p><ul><li>Request enough memory for expected steady operation so scheduling reflects reality.</li><li>Leave measured non-heap and failure headroom below the container limit.</li><li>Test the temporary overlap created by a rolling deployment. Repeat the test while draining a backlog.</li><li>Change one major control at a time and keep the previous resource profile available for rollback.</li></ul><p><strong>Sources: </strong><a href="#reference-kubernetes-resources">Kubernetes</a></p></section>
<section id="native-memory-check"><h2>Compare container memory with JVM memory before raising the limit</h2><p>The output uses placeholders because the values must come from a repeatable workload. Reserved memory is address space the JVM may use later. Committed memory has been made available to the JVM, but it is not identical to the container reading. If committed heap is stable while the container keeps growing, I inspect the NMT categories and operating-system evidence before raising the limit.</p><p>This sequence has strict prerequisites. NMT is off by default and must be enabled when the JVM starts. Oracle documents a performance overhead of roughly 5 to 10 percent, so the choice must be tested for the service. The runtime image must contain a compatible jcmd binary. The command also needs permission to attach to the Java process and must target the correct process ID. NMT does not account for every third-party native allocation. Kubectl top is optimized for recent resource usage rather than precise forensic accounting.</p><div><section><p><small>Step 1</small></p><h3>Read the container memory signal from Kubernetes</h3><p>Start with the memory value Kubernetes exposes for the container. This is a recent operational reading that helps establish whether the container is close to its limit.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>kubectl -n &#x27;&lt;namespace&gt;&#x27; top pod &#x27;&lt;pod&gt;&#x27; --containers</code></pre><p><strong>Sanitized output</strong></p><pre><code>POD    NAME         CPU(cores)    MEMORY(bytes)
&lt;pod&gt;  &lt;container&gt;  &lt;cpu-usage&gt;   &lt;container-memory&gt;</code></pre></figure><p><strong>What this means: </strong>Use this value as a quick comparison point, not as a forensic baseline. Kubectl top is designed for recent resource usage and may not match operating-system tools exactly.</p></section><section><p><small>Step 2</small></p><h3>Ask the JVM where its committed memory is going</h3><p>Next, request an NMT summary from the Java process. This separates committed heap from JVM-managed native areas such as thread stacks and garbage-collector structures.</p><figure><figcaption><strong>Automexia terminal capture</strong></figcaption><pre><code>kubectl -n &#x27;&lt;namespace&gt;&#x27; exec &#x27;&lt;pod&gt;&#x27; -c &#x27;&lt;container&gt;&#x27; -- jcmd &#x27;&lt;java-pid&gt;&#x27; VM.native_memory summary scale=MB</code></pre><p><strong>Sanitized output</strong></p><pre><code>&lt;java-pid&gt;:
Native Memory Tracking:
Total: reserved=&lt;reserved&gt;MB, committed=&lt;committed&gt;MB
- Java Heap (reserved=&lt;heap-reserved&gt;MB, committed=&lt;heap-committed&gt;MB)
- Thread    (reserved=&lt;thread-reserved&gt;MB, committed=&lt;thread-committed&gt;MB)
- GC        (reserved=&lt;gc-reserved&gt;MB, committed=&lt;gc-committed&gt;MB)</code></pre></figure><p><strong>What this means: </strong>If committed heap is stable while container memory keeps growing, inspect the native categories and operating-system evidence. Do not raise the limit until the missing memory has a plausible owner.</p></section></div><p><strong>Sources: </strong><a href="#reference-kubernetes-top">Kubernetes</a>, <a href="#reference-java-nmt">Oracle Java</a></p></section>
<section id="cost-boundary"><h2>Do not claim cloud savings without an auditable cost baseline</h2><p>The previous article promised lower cloud cost but contained no production cost data. I removed that claim. Lower CPU or fewer pods may improve efficiency, but the invoice can still rise because traffic grew or pricing changed. Shared infrastructure also makes one service difficult to price without an allocation rule.</p><p>If cost becomes an objective, I would measure cost per successful payment and per generated document. The calculation must assign a fair share of the cluster and its observability services to each workload. Storage and network transfer must also be included when they materially change. Only then can a capacity change be called a cost improvement rather than an efficiency hypothesis.</p><blockquote>A trustworthy performance report may conclude “faster and more stable; cost impact not measured.” That is stronger than an unsupported savings percentage.</blockquote></section><section><h2>Key takeaways</h2><ul><li>Define the work unit and arrival shape before interpreting resource graphs.</li><li>Separate time spent doing work from time spent waiting in a queue or on a dependency.</li><li>Reactive code still needs a clear boundary around blocking work and scarce dependencies.</li><li>Accept a tuning change only when the workload still recovers safely after failure.</li><li>Do not publish a cloud-cost result without an auditable unit-cost baseline.</li></ul></section><section><h2>References &amp; revisions</h2><ol><li id="reference-spring-webflux"><a href="https://docs.spring.io/spring-framework/reference/web/webflux.html">Spring WebFlux reference</a> — Spring Framework</li><li id="reference-rabbitmq-confirms"><a href="https://www.rabbitmq.com/docs/confirms">Consumer acknowledgements, publisher confirms, and prefetch</a> — RabbitMQ</li><li id="reference-redis-cache"><a href="https://redis.io/docs/latest/develop/reference/client-side-caching/">Client-side caching reference</a> — Redis</li><li id="reference-kubernetes-resources"><a href="https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/">Resource management for Pods and containers</a> — Kubernetes</li><li id="reference-kubernetes-top"><a href="https://kubernetes.io/docs/reference/kubectl/generated/kubectl_top/">kubectl top reference and measurement limits</a> — Kubernetes</li><li id="reference-java-nmt"><a href="https://docs.oracle.com/en/java/javase/21/vm/native-memory-tracking.html">Native Memory Tracking in Java 21</a> — Oracle Java</li></ol><h3>Revision history</h3><p><time datetime="2026-08-25">2026-08-25</time> — Narrowed the article to two real workloads. Container and JVM memory checks now use compact, copyable Automexia captures with colored diagnostic output.</p></section>]]></content:encoded>
      </item>
      </channel>
    </rss>