What it takes to maintain many API integrations at once
An integration is budgeted around the first call and paid for over the next two years. What actually shows up later, with cases we have paid for ourselves.
An integration is not shipped, it is maintained. The budget is built around the first call, the one you see in the demo, and the real cost lands over the next two years, when the provider changes something that was never up to you.
Here is what shows up later, taken from cases we have paid for while keeping integrations with every major social network alive at the same time inside PlanVortex.
No two connection flows are alike, and the budget never says so
The first thing anyone integrating more than one provider finds out is that "connect the user's account" is not one thing. It is four, and only the first one appears in the tutorials:
- The plain OAuth dialog. The user leaves for the provider's site, authorises, and comes back with a code you exchange. This is the easy case and the one nobody argues about.
- A popup that returns nothing in the URL. WhatsApp Business onboarding is a Meta wizard the front end opens with its own JavaScript SDK, and it hands the identifiers back over
postMessage. There is no return parameter to read: if your design assumes every connection ends in a URL with a code, this one has nowhere to live. - A bot you talk to in a chat. Telegram has no authorisation dialog, no code and no account token. The link opens a conversation with a bot, the user drops that bot into their channel, and the account is born minutes later from an incoming webhook, when there is no pending request left to answer.
- The client's own credentials. On Discord the application belongs to the client rather than to us, because the permission that lets an app read message text is reviewed per application once a community gets big enough. With one shared application, the first large client drags everyone else into an annual review cycle.
None of these is an exotic edge case. They are four major providers, and the difference is not in the API but in the moment of connection, which is exactly what gets budgeted as "the connect account screen".
Providers do not always fail with an error
The assumption that costs the most is that a failure arrives as a failure. Three real examples, all of them either HTTP 200 or a message that says the opposite of what happened:
Slack answers 200 with {"ok": false}. For a revoked token, for a channel that does not exist, for a bot that is not in the channel, for a file that is too large. The HTTP client throws nothing, so code that only checks the status code marks the post as sent. The result is a publication recorded as published that exists in no channel.
Discord answers 200 with empty text. If the application is missing the message content permission, the API does not return a 403: it returns the list of messages with the text field blank. Store that and the problem stops being an error and becomes a database full of empty comments that look real.
Telegram answers 400 for two opposite things. "message to delete not found" means the message is gone, which is where we wanted to get. "message can't be deleted" means it is still published and the user needs to hear about it. The only thing separating success from failure is the sentence.
Hence a rule that is not about style: all traffic to a provider goes through a single place, and that is where what counts as an error gets decided. One stray call anywhere else in the codebase skips that gate, and with it every bit of that interpretation.
The rate limit is not yours, it is everyone's
Providers do not count requests per customer of yours. They count per application or per IP address, which turns one customer's bug into an outage for all of them.
Discord shows this best. It counts invalid requests (401, 403, 429) separately, and its network provider blocks by IP at 10,000 of them in ten minutes. That is not an application ban: it is your whole server, with all your customers inside it, and giving each customer their own application does not help, because the IP is the same.
Against that, a statistical strategy (backoff, a bit of jitter) is not enough, because it promises average behaviour and the ban is caused by the worst ten minutes of the month. What works is an arithmetic guarantee: one shared token bucket for the whole process, with a ceiling in requests per second the code cannot exceed. At ten per second, a bug failing 100% of its requests tops out at 6,000 invalid requests in ten minutes, below the ban. That sentence can go in a contract; "we retry with exponential backoff" cannot.
And there is a worse case, the genuinely shared resource: when the bot is one for the entire platform (Telegram), the provider's limit belongs to all customers at once. A noisy customer does not throttle themselves, they throttle everyone else. The queue has to be per conversation and the ceiling global.
Mocked tests stay green forever
A test that compares your request against a response you wrote yourself checks what your code does, not what the API accepts. It is useful (it is what tells you that you broke the parsing), but its blind spot is enormous: it passes just as happily the day the provider retires a version, requires a new field, or stops returning a value.
The most expensive case we have had was not even a rejection. When uploading an image to Slack, the step that creates the message does not return the identifier of the message it created, however natural that would seem. Our mocked test had been returning it for months, green, because whoever wrote it assumed the reasonable thing. In production, every publication with an image ended up flagged as failed. The API rejected nothing: it simply answered with less than the mock took for granted.
That is why the integrations we maintain have three layers, not two:
| Layer | What it checks | What it cannot see |
|---|---|---|
| Our own logic | The state machine, retries, isolation between accounts | Anything happening on the other side |
| Mocked contract | Which request gets built and how the response is read | That the request is no longer accepted |
| Real API | That the provider still behaves the way we think | Nothing: it is the only one that can say |
The third one costs money and time (some APIs charge per call), so it runs separately and writes nothing by default. It is also the only one that finds what matters.
The part of the schedule you do not control
Almost every large integration has a review step before production: Meta, Google and TikTok review the application and every permission it requests. It is measured in weeks and cannot be compressed by adding people.
What you can do is not lose the first round, and that is prepared from the proposal onwards: the app genuinely working so you can record the video they ask for, the privacy policy published on a verified domain, and every permission justified with the product screen where it is used. A permission requested "just in case" is the most common rejection reason we have run into.
It belongs in the schedule as an external dependency, like a hardware supplier's lead time. If it turns up the week before launch, the launch date no longer exists.
Underneath it: an honest interface
With several providers at once, the temptation is to hide them behind one common interface that promises the same thing for all of them. It works until you look closely: some providers do not publish at all (a local listing receives reviews, not posts), some will not let you delete a third party's comment, some have no private messages, and some never notify you of anything so you have to go and ask.
What has held up for us is the opposite: a common interface for what genuinely is common, and declared capabilities for the rest. Each provider states what it can do, and the product (the dashboard, the public API, the customer integrating it) reads that and behaves accordingly instead of trying and failing. Showing a button the provider does not support is a bug only the end user gets to see.
What to ask before signing off on an integration
Four questions that change the budget, and none of them is about the first call:
- Who watches the failures, and how often? An integration with nobody watching is an integration you find out is broken from a customer.
- Are there tests that talk to the real API? If not, the first notice of a change will be an incident.
- What is the request ceiling, and who shares it? That is the difference between one affected customer and all of them.
- What happens the day the permission expires? Which is a big enough problem to have its own article.
We budget integrations with all of this inside rather than as an add-on: it is part of what the integrations and APIs service includes. If what you need is someone to keep it running after delivery, that is exactly the part that has to be written into the contract.
Frequently asked questions
- How much does it cost to maintain an API integration per year?
- It depends on how much the provider changes, and that is not your call. What you can budget is the capacity around it: who watches the failures, how often the tests that talk to the real API run, and what happens when the provider retires an endpoint with three months' notice. An integration shipped without that agreement works until the first change on the other side, and then the cost shows up anyway, as an emergency.
- Why aren't mocked tests enough for an integration?
- Because a test built on a response you wrote yourself checks what your code does, not what the API accepts. It stays green the day the provider requires a new field, retires a version, or stops returning a value your code assumed. You also need a suite that talks to the real API, even if it runs weekly and only reads.
- What delays an integration the most?
- The provider's review. Meta, Google and TikTok review the application and every permission it asks for before letting it into production, and that is measured in weeks, not days. You prepare for it from the proposal onwards (a video of the app actually working, a published privacy policy, a verified domain), because losing the first round costs more than doing it properly.
- Can an integration layer be provider-agnostic?
- Partly. The inside of it (queues, retries, permission renewal, rate limiting) is yours and works for everyone. What cannot be abstracted away are the real differences: some providers do not publish at all, some will not let you delete a third party's comment, some have no private messages. Those get declared and surfaced, not hidden behind an interface that promises the same thing for all of them.