The hardest part of a multi-cloud migration is keeping the message path clear while services and infrastructure move at different speeds.
Before moving traffic, I answer three questions. Which path is active? What will prove that a message completed? What exact action returns traffic if the new path fails? Those answers turn a risky migration into a sequence the team can observe and reverse.
The production problem is an ambiguous message path
A green health check proves that the connector process is running. It does not prove that the expected message reached the correct consumer. A pod can be ready while its consumer is attached to the wrong topic, so process health and flow health must be checked separately.
I follow one known message from the producer to the business outcome. An opaque correlation ID travels with the message and lets an operator find it without searching for customer data. At each boundary, I ask whether the message entered and whether it left. The first boundary without evidence becomes the starting point for the investigation. I also record which team owns that boundary, so the incident does not stall while people search for the right operator.
The decision I changed was where recovery ends. Connector health is still useful, but it is too narrow to be the final checkpoint. I now end the evidence path at the business outcome. This prevents a healthy process from hiding work that is delayed, duplicated, or missing.
| Checkpoint | Evidence to capture | What it rules out |
|---|---|---|
| Producer | Find the message by its correlation ID and confirm the destination. | Shows whether the application emitted the message. |
| Broker boundary | Compare the source position with the replicated position. | Shows whether the migration path moved the message. |
| Consumer | Track the known message while confirming that the consumer offset continues to advance. The offset is the broker position saved for that consumer group. | Shows whether a healthy process is actually consuming. |
| Business outcome | Confirm the terminal event or locate the item in a reconciliation report. | Prevents transport success from being mistaken for user success. |
During a migration, “the connector is up” is an infrastructure statement. “The expected business event completed once” is a recovery statement.
Write the cutover contract before touching production
The most useful migration document I keep is one page for one flow. The first line names the path that currently owns the traffic and the path that will replace it. The next line names the person who can approve the move and how long the team will observe it.

The record then defines proof. For example, it may require the consumer offset to keep advancing and a business completion event to appear. It also defines one stop signal and one tested route back. The thresholds stay as placeholders until the service owner chooses values that fit the workload.
This removes a dangerous debate from the incident channel. If the new consumer falls behind before cutover, the old path remains authoritative while the team investigates. If traffic already moved and message correctness becomes uncertain, the team follows the tested return route. A human still makes the decision, but the evidence and safe options are already clear.
flow: <business-flow>
current_source_of_truth: on_prem
target: azure
owner: <team-and-on-call>
observe:
- producer_send_rate
- mirror_lag
- consumer_offset_progress
- business_completion_rate
abort_when:
mirror_lag: <agreed-workload-limit>
duplicate_rate: <agreed-workload-limit>
completion_rate: <agreed-slo-limit>
reverse:
route: <tested-return-path>
data_check: <reconciliation-query>SourcesApache Kafka
Use the first fifteen minutes to remove ambiguity
When a migrated flow degrades, I do not start by blaming a cloud platform or broker. I first name the exact business flow and the direction in which it is failing. Then I record when the problem began and the last outcome we know succeeded. This gives every team the same incident to investigate.

The sequence below is a response drill derived from the migration work. It is intentionally evidence-first. Root cause can wait; an incorrect traffic move can make reconciliation much harder.
- Incident commander: sets the next objective and decides whether a production change is allowed.
- Flow lead: follows one message path end to end and prevents unrelated debugging branches.
- Platform owners: inspect their boundary and return timestamped evidence instead of saying that it looks healthy.
- Scribe: records each query and production change with a timestamp. This makes later reconciliation possible.
| Window | Coordinator | Technical lead |
|---|---|---|
| 0-3 minutes | State which flow is failing. Assign one person who can approve a change. | Find the last message that reached its expected business outcome. |
| 3-7 minutes | Pause unrelated changes. Keep one shared timeline for the response. | Trace one message forward and mark the first checkpoint where evidence disappears. |
| 7-12 minutes | Ask for one reversible mitigation and the signal that would stop it. | Apply the smallest safe action to the affected path. |
| 12-15 minutes | Publish one short update. State the next action and the time of the next update. | Confirm that the business outcome recovered, not merely the pod. |
Separate process health from flow health
A readiness probe answers whether the connector should receive work. A liveness probe answers whether restarting it could restore progress. Those are different decisions. If liveness restarts a connector that is merely slow, the restart can trigger a consumer rebalance. During a rebalance, Kafka redistributes partitions between consumer instances, so useful processing may pause and the backlog may grow.
I use Dynatrace to find when the error began and which downstream call changed. I use Conduktor to see whether the consumer offset is still moving. Then I connect both views with one correlation ID and confirm the final business state. Neither tool can prove completion on its own.
- A flat producer rate with falling completion suggests the break is downstream of emission.
- Growing lag with stable consumer membership points toward processing capacity or a blocked dependency.
- Frequent group rebalances after pod restarts make probe configuration and shutdown behavior suspects.
- Stable offsets with missing business acknowledgements move the investigation into processing or persistence.
SourcesKubernetesDynatrace
Collect read-only evidence before changing the flow
The sequence below is a sanitized reconstruction, not a production transcript. It narrows the failed boundary without changing traffic or message positions.
Every command requires approved read access. Search logs only with an opaque correlation ID. Keep Kafka credentials inside a protected client configuration instead of placing them on the command line.
Check whether Kubernetes has been restarting the connector
Start with process stability. This command shows the connector pods and their restart counts. It is a fast way to see whether the platform has been repeatedly replacing the process.
kubectl -n '<namespace>' get pods -l 'app=<connector>'| NAME | READY | STATUS | RESTARTS |
|---|---|---|---|
| <connector-pod> | 1/1 | Running | <count> |
Find the known message in recent connector logs
Next, search for one approved correlation ID. The goal is to confirm whether this connector recorded the message and which topic it attempted to use.
kubectl -n '<namespace>' logs 'deployment/<connector>' --since=15m --prefix | grep -- '<correlation-id>'<pod> <timestamp> correlation_id=<correlation-id> published topic=<topic>Check whether the Kafka consumer group is advancing
Finally, inspect the consumer-group position. CURRENT-OFFSET is the position saved by the consumer. LOG-END-OFFSET is the newest position in the partition. LAG is the difference between them.
bin/kafka-consumer-groups.sh --bootstrap-server '<broker>' --command-config '<protected-client.properties>' --describe --group '<consumer-group>'| TOPIC | PARTITION | CURRENT-OFFSET | LOG-END-OFFSET | LAG | CONSUMER-ID |
|---|---|---|---|---|---|
| <topic> | <partition> | <committed> | <latest> | <difference> | <consumer-id> |
SourcesKubernetesApache Kafka
Recovery includes reconciliation and a tested way back
A migration incident is not over when lag returns to zero. The consumer may have processed a message twice, or the producer may have skipped it before the broker ever saw it. I close the gap with a business reconciliation that finds missing and duplicate outcomes without exposing sensitive payloads.
Failback needs the same care as failover. First, confirm which environment is allowed to write. Next, prove that its checkpoint maps to the message position you expect. Keep mirroring active until the recovered path has handled normal and retry traffic for the agreed observation window. Removing it earlier can erase the safest route back.
| Dimension | Recovery question | Evidence |
|---|---|---|
| Traffic | Is expected work entering the authoritative path? | Producer and ingress rates by flow |
| Progress | Is the oldest queued item getting newer rather than older? | Watch queue age while confirming that offsets move. |
| Correctness | Did work complete once and in the right state? | Domain reconciliation |
| Capacity | Can the recovered path absorb retries without falling behind again? | Drain a controlled backlog and watch saturation. |
| Reversal | Can we still return safely if the symptom recurs? | Run the route check and confirm which side owns writes. |
What this experience changed in my practice
The migration reinforced that cross-environment reliability is mostly an ownership and evidence problem. Mirroring creates a return path, but it also creates more states for responders to understand. I now keep a one-page record for each important flow. The first line says which path is active and who owns it. The next line contains the query that proves the business outcome.
Private thresholds and incident results are intentionally omitted, so this article does not claim an improvement in mean time to recovery. The useful result is the operating model. A future lab can demonstrate how lag and redelivery behave during a safe reversal with synthetic data.
- Name the authoritative path for each stage of the migration.
- Define recovery at the business boundary, not at the process boundary.
- Use transport evidence to locate the break. Use business reconciliation to prove recovery.
- Test failback while the original path is still healthy enough to teach you.
Key takeaways
- A live cloud migration needs an evidence map for every critical message flow.
- The first response objective is to name the affected flow and find the first boundary without evidence.
- A healthy pod does not prove that a message completed. Verify the business outcome separately.
- Mirroring is useful only when the team has tested the return route and knows which environment may write.
- Honest anonymization is more credible than invented incident metrics.
References & revisions
Primary documentation supports the technical guidance below.
- MirrorMaker and geo-replication documentation Apache Kafka
- Liveness, readiness, and startup probes Kubernetes
- Distributed tracing concepts and analysis Dynatrace
- Cloud Adoption Framework migration guidance Microsoft Azure
- kubectl logs reference Kubernetes
- Basic Kafka operations and consumer-group inspection Apache Kafka
Revision history
Rebuilt the article around a step-by-step evidence path. Its read-only commands now use compact, copyable Automexia captures with separate sanitized output.



Ubuntu-24.04