The hardest part of a multi-cloud migration is keeping the message path clear while services and infrastructure move at different speeds.

Before moving traffic, I answer three questions. Which path is active? What will prove that a message completed? What exact action returns traffic if the new path fails? Those answers turn a risky migration into a sequence the team can observe and reverse.

The production problem is an ambiguous message path

A green health check proves that the connector process is running. It does not prove that the expected message reached the correct consumer. A pod can be ready while its consumer is attached to the wrong topic, so process health and flow health must be checked separately.

I follow one known message from the producer to the business outcome. An opaque correlation ID travels with the message and lets an operator find it without searching for customer data. At each boundary, I ask whether the message entered and whether it left. The first boundary without evidence becomes the starting point for the investigation. I also record which team owns that boundary, so the incident does not stall while people search for the right operator.

The decision I changed was where recovery ends. Connector health is still useful, but it is too narrow to be the final checkpoint. I now end the evidence path at the business outcome. This prevents a healthy process from hiding work that is delayed, duplicated, or missing.

Evidence checkpoints for a migrating flowThe control is useful only when it proves progress at the next boundary, not merely local health.
CheckpointEvidence to captureWhat it rules out
ProducerFind the message by its correlation ID and confirm the destination.Shows whether the application emitted the message.
Broker boundaryCompare the source position with the replicated position.Shows whether the migration path moved the message.
ConsumerTrack the known message while confirming that the consumer offset continues to advance. The offset is the broker position saved for that consumer group.Shows whether a healthy process is actually consuming.
Business outcomeConfirm the terminal event or locate the item in a reconciliation report.Prevents transport success from being mistaken for user success.
During a migration, “the connector is up” is an infrastructure statement. “The expected business event completed once” is a recovery statement.

Write the cutover contract before touching production

The most useful migration document I keep is one page for one flow. The first line names the path that currently owns the traffic and the path that will replace it. The next line names the person who can approve the move and how long the team will observe it.

An operator validates each step before moving a flow between cloud environments.
A cutover is ready when the new path works and the return path has been tested.

The record then defines proof. For example, it may require the consumer offset to keep advancing and a business completion event to appear. It also defines one stop signal and one tested route back. The thresholds stay as placeholders until the service owner chooses values that fit the workload.

This removes a dangerous debate from the incident channel. If the new consumer falls behind before cutover, the old path remains authoritative while the team investigates. If traffic already moved and message correctness becomes uncertain, the team follows the tested return route. A human still makes the decision, but the evidence and safe options are already clear.

AutomexiaYAML
Sanitized cutover decision record
flow: <business-flow>
current_source_of_truth: on_prem
target: azure
owner: <team-and-on-call>
observe:
  - producer_send_rate
  - mirror_lag
  - consumer_offset_progress
  - business_completion_rate
abort_when:
  mirror_lag: <agreed-workload-limit>
  duplicate_rate: <agreed-workload-limit>
  completion_rate: <agreed-slo-limit>
reverse:
  route: <tested-return-path>
  data_check: <reconciliation-query>
This is a sanitized template. Every placeholder must be agreed and tested for the specific flow.

SourcesApache Kafka

Use the first fifteen minutes to remove ambiguity

When a migrated flow degrades, I do not start by blaming a cloud platform or broker. I first name the exact business flow and the direction in which it is failing. Then I record when the problem began and the last outcome we know succeeded. This gives every team the same incident to investigate.

An incident commander assigns clear investigation and communication responsibilities.
Coordination and investigation need separate owners when several platforms are involved.

The sequence below is a response drill derived from the migration work. It is intentionally evidence-first. Root cause can wait; an incorrect traffic move can make reconciliation much harder.

  • Incident commander: sets the next objective and decides whether a production change is allowed.
  • Flow lead: follows one message path end to end and prevents unrelated debugging branches.
  • Platform owners: inspect their boundary and return timestamped evidence instead of saying that it looks healthy.
  • Scribe: records each query and production change with a timestamp. This makes later reconciliation possible.
First-fifteen-minute response drillTimes are coordination targets for the drill, not reported timings from a specific outage.
WindowCoordinatorTechnical lead
0-3 minutesState which flow is failing. Assign one person who can approve a change.Find the last message that reached its expected business outcome.
3-7 minutesPause unrelated changes. Keep one shared timeline for the response.Trace one message forward and mark the first checkpoint where evidence disappears.
7-12 minutesAsk for one reversible mitigation and the signal that would stop it.Apply the smallest safe action to the affected path.
12-15 minutesPublish one short update. State the next action and the time of the next update.Confirm that the business outcome recovered, not merely the pod.

Separate process health from flow health

A readiness probe answers whether the connector should receive work. A liveness probe answers whether restarting it could restore progress. Those are different decisions. If liveness restarts a connector that is merely slow, the restart can trigger a consumer rebalance. During a rebalance, Kafka redistributes partitions between consumer instances, so useful processing may pause and the backlog may grow.

I use Dynatrace to find when the error began and which downstream call changed. I use Conduktor to see whether the consumer offset is still moving. Then I connect both views with one correlation ID and confirm the final business state. Neither tool can prove completion on its own.

  • A flat producer rate with falling completion suggests the break is downstream of emission.
  • Growing lag with stable consumer membership points toward processing capacity or a blocked dependency.
  • Frequent group rebalances after pod restarts make probe configuration and shutdown behavior suspects.
  • Stable offsets with missing business acknowledgements move the investigation into processing or persistence.

SourcesKubernetesDynatrace

Collect read-only evidence before changing the flow

The sequence below is a sanitized reconstruction, not a production transcript. It narrows the failed boundary without changing traffic or message positions.

Every command requires approved read access. Search logs only with an opaque correlation ID. Keep Kafka credentials inside a protected client configuration instead of placing them on the command line.

Step 01

Check whether Kubernetes has been restarting the connector

Start with process stability. This command shows the connector pods and their restart counts. It is a fast way to see whether the platform has been repeatedly replacing the process.

AutomexiaUbuntu-24.04
kubectl -n '<namespace>' get pods -l 'app=<connector>'
Sanitized output✓  184ms
NAMEREADYSTATUSRESTARTS
<connector-pod>1/1Running<count>
Step 02

Find the known message in recent connector logs

Next, search for one approved correlation ID. The goal is to confirm whether this connector recorded the message and which topic it attempted to use.

AutomexiaUbuntu-24.04
kubectl -n '<namespace>' logs 'deployment/<connector>' --since=15m --prefix | grep -- '<correlation-id>'
Sanitized output✓  312ms
<pod> <timestamp> correlation_id=<correlation-id> published topic=<topic>
Step 03

Check whether the Kafka consumer group is advancing

Finally, inspect the consumer-group position. CURRENT-OFFSET is the position saved by the consumer. LOG-END-OFFSET is the newest position in the partition. LAG is the difference between them.

AutomexiaUbuntu-24.04
bin/kafka-consumer-groups.sh --bootstrap-server '<broker>' --command-config '<protected-client.properties>' --describe --group '<consumer-group>'
Sanitized output✓  428ms
TOPICPARTITIONCURRENT-OFFSETLOG-END-OFFSETLAGCONSUMER-ID
<topic><partition><committed><latest><difference><consumer-id>

SourcesKubernetesApache Kafka

Recovery includes reconciliation and a tested way back

A migration incident is not over when lag returns to zero. The consumer may have processed a message twice, or the producer may have skipped it before the broker ever saw it. I close the gap with a business reconciliation that finds missing and duplicate outcomes without exposing sensitive payloads.

Failback needs the same care as failover. First, confirm which environment is allowed to write. Next, prove that its checkpoint maps to the message position you expect. Keep mirroring active until the recovered path has handled normal and retry traffic for the agreed observation window. Removing it earlier can erase the safest route back.

Exit conditions before closing the incidentThe exact thresholds belong to the service owner and its service-level objective; these dimensions should not be skipped.
DimensionRecovery questionEvidence
TrafficIs expected work entering the authoritative path?Producer and ingress rates by flow
ProgressIs the oldest queued item getting newer rather than older?Watch queue age while confirming that offsets move.
CorrectnessDid work complete once and in the right state?Domain reconciliation
CapacityCan the recovered path absorb retries without falling behind again?Drain a controlled backlog and watch saturation.
ReversalCan we still return safely if the symptom recurs?Run the route check and confirm which side owns writes.

What this experience changed in my practice

The migration reinforced that cross-environment reliability is mostly an ownership and evidence problem. Mirroring creates a return path, but it also creates more states for responders to understand. I now keep a one-page record for each important flow. The first line says which path is active and who owns it. The next line contains the query that proves the business outcome.

Private thresholds and incident results are intentionally omitted, so this article does not claim an improvement in mean time to recovery. The useful result is the operating model. A future lab can demonstrate how lag and redelivery behave during a safe reversal with synthetic data.

  • Name the authoritative path for each stage of the migration.
  • Define recovery at the business boundary, not at the process boundary.
  • Use transport evidence to locate the break. Use business reconciliation to prove recovery.
  • Test failback while the original path is still healthy enough to teach you.

Key takeaways

  • A live cloud migration needs an evidence map for every critical message flow.
  • The first response objective is to name the affected flow and find the first boundary without evidence.
  • A healthy pod does not prove that a message completed. Verify the business outcome separately.
  • Mirroring is useful only when the team has tested the return route and knows which environment may write.
  • Honest anonymization is more credible than invented incident metrics.

References & revisions

Primary documentation supports the technical guidance below.

  1. MirrorMaker and geo-replication documentation Apache Kafka
  2. Liveness, readiness, and startup probes Kubernetes
  3. Distributed tracing concepts and analysis Dynatrace
  4. Cloud Adoption Framework migration guidance Microsoft Azure
  5. kubectl logs reference Kubernetes
  6. Basic Kafka operations and consumer-group inspection Apache Kafka

Revision history

Rebuilt the article around a step-by-step evidence path. Its read-only commands now use compact, copyable Automexia captures with separate sanitized output.