Investigate integration incidents
Integration failures are usually reconstructed across application logs, provider dashboards, support tickets, and release history. Each source holds part of the story; none survives as an investigation record. Seamward groups the related evidence into one incident whose every claim links to something observed.
The scenario
A customer reports that some new candidates are missing salary data. Support escalates. Without a shared record, the on-call engineer starts from zero: grep the logs for the customer's requests, check the provider's status page (green), diff recent deploys, ask in Slack whether anyone changed the mapping. An hour in, they are still assembling timestamps.
With the integration observed, the same engineer opens one incident instead. At its head sits the finding that started it, with the exact path:
Code example
{ "kind": "type-changed", "path": "$.candidate.salary", "expected": "string", "observed": "number"}One timeline instead of five tools
The incident already contains what the first hour of investigation normally produces, in order:
Code example
09:24 First observation with changed shape release 2026.08.12, commit 9f831dq09:24 Finding: type-changed at $.candidate.salary expected string, observed number09:31 14 further observations share the new shape same release still serving10:02 Incident opened: findings grouped first/latest occurrence recorded10:40 Replay dataset built from incident evidence immutable, 15 fixtures10:52 Repair proposal generated and validated replay: pass11:15 Approved by dana@acme.com rationale recordedThe release context row is the hour-saver: the shape changed while a twelve-day-old release was serving steadily, so this is provider drift, not a bad deploy, and nobody needs to bisect their own history to prove it.
The incident boundary
An incident stays scoped to related evidence from one integration and environment. A narrow boundary is what keeps the timeline reviewable:
- Several observations reporting the same structural change: one incident.
- An unrelated provider failing the same afternoon: a different incident, even if the same customer noticed both.
- A missing business outcome needing its triggering request and release context: one incident holding both.
Join evidence only when correlation and timing support one investigation; split everything else.
Record the conclusion
The timeline separates three things that reconstructions blur: what was observed (the findings and observations), what was inferred (the explanation the team settled on), and what was decided (the reviewer's approval or rejection, with provenance). Generated summaries stay advisory and never override linked evidence. Unresolved questions stay visible: a useful incident record makes uncertainty explicit instead of promoting a plausible hypothesis to a fact.
Resolving an incident records workflow state, a reason, a timestamp, and an audit event. It does not delete evidence, and it does not claim a fix was deployed. Missing-outcome incidents verify sustained recovery for at least five minutes, or the rule's longer delay window, before auto-resolving with reason recovered. An administrator or owner may instead record an accepted risk, false positive, or superseded incident with a required note and attributed reviewer snapshot.
For a merged pull-request repair, the first observation carrying the merge commit establishes the deployment boundary. Seamward requires healthy correlated traffic after that boundary before verification starts, preserves the original failure evidence, and resets verification if a new missing outcome appears.
Operating boundaries
- Workspace scope is enforced for every incident and linked record.
- Generated summaries do not override observations or deterministic findings.
- Incident status records workflow state; it does not prove remediation succeeded.
- Repair generation and delivery remain separate, review-gated actions.
Related: Incidents for boundary and status semantics, and Investigate an incident for the dashboard workflow.
