Connecting two systems can remove a surprising amount of repetitive work. A new member signs up on your website and appears in your CRM. A donation reaches your payment provider and is recorded in your accounting system. A registration updates an email list without anyone copying a row from one spreadsheet to another.
When it works, the integration almost disappears. That is precisely why it can be neglected.
Before you connect two systems, decide what should happen when the connection fails. Identify the task the integration supports, which system holds the authoritative record, how failures will be detected, who will respond, and how the work can continue safely in the meantime.
That may sound less exciting than drawing arrows between colourful boxes. It is also what turns a clever connection into a dependable service.
Start with the work, not the software
An integration diagram might say:
Website → CRM → email platform
The people involved experience something more concrete:
A person registers for a program, receives confirmation, and expects the organization to know that they are coming.
That second description gives you a much better starting point. It identifies an outcome that matters to people and makes the consequences of failure easier to see.
If the website records the registration but the CRM does not, can staff still find the person? If the email platform is delayed, will the person assume their registration failed and submit it again? If two copies arrive later, could they be counted twice?
Technical reliability should be defined in terms of the user’s experience. Google Cloud’s reliability guidance makes the same point more formally: reliability goals should be grounded in user experience, with systems designed to detect problems, respond, recover, and learn from failures.
Before discussing APIs, write down:
- the real-world task the integration supports;
- what the user believes has happened at each step;
- what staff need in order to complete the work;
- the harm caused by delay, duplication, or lost information.
This is not preamble. It determines which technical safeguards are worth building.
Choose which system tells the truth
Once the same information exists in two places, a simple question becomes important: which copy is authoritative?
Suppose a member updates their email address in your online account, while a staff member updates a different address in the CRM a few minutes later. Which change wins? Does information move in one direction or both? Can an old value overwrite a newer one?
For each important piece of information, name a system of record. You might decide that:
- the payment provider is authoritative for whether a payment succeeded;
- the membership database is authoritative for membership status;
- the website account is authoritative for a user’s communication preferences;
- the accounting platform is authoritative for financial categorization.
There may be good reasons to make different choices. What matters is that the choice is explicit. Without it, a synchronization problem can become a quiet disagreement between systems, with staff left to decide which screen they trust.
Also define direction and timing. Does data move immediately, every fifteen minutes, or overnight? A five-minute delay may be irrelevant for a mailing address and unacceptable for an access-control decision.
Design for partial success
Integrations rarely fail as neatly as a light switching off. One step may succeed while the next does not.
A payment can be completed even if the CRM update times out. A person can be added to the membership database while the confirmation email remains unsent. Retrying the whole sequence carelessly may then create a second payment, record, or message.
Developers use the word idempotency for an operation that can safely be repeated without repeating its effect. Stripe supports idempotency keys so a request can be retried after a connection problem without accidentally creating the same transaction twice. AWS likewise recommends idempotent operations because network failures and duplicate message delivery are normal possibilities in distributed systems.
You do not need to use that terminology in a planning meeting. You do need to ask the practical question:
If we are unsure whether this step worked, can we try it again safely?
For any action involving money, enrolment, inventory, permissions, or outbound communication, the answer deserves careful attention.
Retries also need limits. Repeating a request immediately and indefinitely can make an overloaded service worse. AWS recommends delaying repeated attempts and distinguishing temporary failures from errors that will not be fixed by trying again. Your technical team can choose the mechanism. Your organization should still decide how long a delay is acceptable and when a person needs to take over.
Make failure visible to the right person
A failed integration should not depend on a customer noticing first.
Logging an error is useful, but only if someone knows the log exists and is expected to look at it. An effective alert should answer enough questions to begin a response:
- Which connection failed?
- When did it fail?
- Which records or people may be affected?
- Will the system retry automatically?
- What action, if any, should staff take?
Avoid sending every technical hiccup to a crowded shared inbox. If the software recovers automatically within a minute and no one is affected, a record for later review may be enough. If a donation was accepted but will not appear in the donor database, the appropriate staff member may need to know promptly.
The alerting rule should follow the consequence, not merely the error code.
Give staff a safe recovery path
When an integration is unavailable, “do it manually” sounds reassuringly simple. Manual work can also create a second source of confusion if nobody defines how it should happen.
A useful fallback procedure explains:
- where staff can find the original record;
- how they can tell whether it has already reached the destination;
- which fields may be entered or changed manually;
- how the manual action is marked to prevent later duplication;
- what should happen when the connection returns.
Test this with someone who did not build the integration. If the recovery process requires a developer to query a database and interpret several identifiers, you do not yet have a routine operational fallback. You have an escalation path—and that may be acceptable, provided it is documented and the right person is available.
Name an owner on both sides of the connection
Every integration needs technical care, but it also needs operational ownership.
The technical owner understands authentication, data mapping, logs, retries, and changes to the connected services. The operational owner understands what the information means, how quickly it is needed, and what staff or users should be told when it is delayed.
For a small team, one person may cover both roles. The distinction still helps. It prevents an integration from becoming “the website’s problem” when the real questions concern finance, membership, communications, or program delivery.
Record at least:
- the purpose of the integration;
- the systems and vendors involved;
- the technical and operational owners;
- where credentials are managed (not the credentials themselves);
- the system of record for each critical field;
- expected timing and acceptable delay;
- how failures are detected;
- the retry and duplicate-prevention approach;
- the manual fallback and reconciliation process;
- vendor renewal dates, usage limits, and relevant support contacts.
This does not need to become a sixty-page manual. A short document that is current, findable, and tested is far more valuable than an exhaustive one that nobody trusts.
Treat vendor changes as part of maintenance
An integration connects systems that continue to evolve independently. A provider may change an API version, permission model, pricing tier, rate limit, or authentication method. A field that was optional may become required. Your own team may change a workflow without realizing another system relies on it.
Include integrations in ordinary maintenance rather than waiting for an emergency. Review notices from vendors. Keep a test environment where practical. Check that alerts still reach an active person. Periodically trace one representative record from beginning to end.
Most importantly, revisit the original purpose. A connection that once saved time may now preserve a workflow your team no longer needs. Maintenance includes retiring unnecessary complexity, not only keeping it alive.
Try a failure walkthrough before you build
Choose one proposed or existing integration and gather the people who understand the user task, the operational process, and the technology. Put one example record on the table—a registration, donation, order, referral, or member update—and trace its journey.
At each handoff, ask:
- What confirms that this step succeeded?
- What would happen if it were delayed?
- Could repeating it cause harm?
- Who would notice the problem?
- How would we recover?
You may discover that the integration needs a queue, an alert, a reconciliation report, a clearer owner, or a simpler workflow. You may even decide that a small amount of well-designed manual work is safer than another automated connection.
Either result is useful. The goal is not to connect as many systems as possible. It is to help people complete important work with fewer surprises—and to ensure your team can respond calmly when the invisible connection becomes visible again.
Systems are imperfect, after all, just like the humans who build them. If you plan for failure, it'll sting less when it happens.