OAuth token refresh without disconnecting accounts

Almost no permission lasts forever and the one that expires warns nobody. How token renewal is designed so it does not disconnect accounts, with real clocks.

By Francisco José Fernández-Medina LópezPublished on 8 min
60 d 24 h 60 min ∞

Almost no permission lasts forever, and the one that expires warns nobody. There is no advance notice from the provider, no email, no callback: the next call simply fails. And the next call is usually a background job at three in the morning, with nobody watching.

This is what you learn maintaining permission renewal for every major social network at once, and it is the part of an integration that gets budgeted worst, because none of it shows up in the demo.

Every provider has its own clock, and they do not match

The first assumption to throw away is that there is such a thing as "the expiry time". These all coexist, all in production, all at the same time:

Access token lifetime Providers
60 days Meta (Facebook, Instagram), LinkedIn, Threads
24 hours TikTok
60 minutes Google (YouTube, Google Business)
Under 30 minutes Bluesky
Never expires Discord webhook, Slack bot token, Telegram (no account token at all)

And lifetime is not the only thing that varies. What you exchange varies too: some providers return a brand new refresh token every time, some hand it over once and for good, and at least one has no refresh token whatsoever.

Which leads to the first design decision: there is no such thing as "the token refresh job". There are two paths, and they coexist.

A scheduled sweep is not enough

The obvious approach is a job that picks up accounts expiring soon every hour and renews them. That works with 60-day tokens. With 60-minute tokens, an account can expire between two passes, and the job that publishes runs every minute, not every hour.

So you also need the other path: renew on demand, right before using the token, when it has little life left. And then both paths coexist and have to be idempotent, so that whichever gets there first leaves a fresh token and the other does nothing.

Even so the sweep does not become redundant, and the reason is the one people miss: an account nobody uses for weeks never goes through the on-demand path at all. Without a sweep, the session dies quietly without a single call failing, because there are no calls.

The extreme version of this is a provider with a 60-day token and no recovery path: an account connected and left alone for 61 days dies on its own, and reconnecting it is the user's job. There the sweep is not an optimisation, it is the only thing keeping the account alive.

The single-use refresh token

In some protocols the refresh token rotates: redeeming it gives you a new one and burns the old. It is safer, and it breaks the naive design, because it turns an OAuth problem into a concurrency problem.

If the scheduled job and the application refresh the same account at the same time, one of them redeems a token that is already spent. The provider rejects it, and since that response is indistinguishable from a revoked permission, the account ends up flagged as broken. The account disconnects itself, and nobody did anything wrong.

Two things fix it, and you need both:

  • A per-account lock. Only one process renews that account at a time; the other waits and finds the token already fresh. The scheduled job does not redeem anything on its own: it asks for renewal through the same path as everyone else.
  • A write conditioned on the token you started from. When saving the result, require that the token you redeemed is still the one in the database. If it is not, someone else rotated it in the meantime and what is stored is newer, so you leave it alone. Without that condition, a slow process writes a spent token back and the account dies on the next pass.

The second one is the one people skip, because the lock looks sufficient. It is not, when the other process has been holding a stale copy in memory since before the lock existed.

Never overwrite what the provider did not send you

This is a one-line bug that kills accounts by the network.

The normal case is a refresh response carrying a new access token, a new refresh token and both expiry times, so you assign all three. And then along comes a provider that returns no refresh token, either because it handed it over once at connection time and it does not expire (Google), or because it does not exist and what you present in order to refresh is the access token itself (Threads).

Assigning the response blindly leaves the field empty. The account keeps working until the next renewal, which now has nothing to renew with. In practice, the first pass of the scheduled job kills every account on that network, and the customer finds them disconnected with no explanation.

The rule is short: only overwrite what the provider actually returned. And if the language allows it, make the field optional in the response type, so the compiler is what complains the day somebody removes the check.

A transient failure must not disconnect anyone

When a renewal fails you have to decide whether to flag the account, and both decisions cost money if you get them wrong.

Flagging it takes it out of the automatic cycle, which is exactly what you want for a revoked permission: retrying hourly against a token that no longer works fixes nothing and burns quota. But if the failure was a timeout or a 500 from the provider, that account sits there in red until the user reconnects something that was never broken.

So before flagging you separate the transient failure from the permanent one, and you store the specific code, not one generic error for everything. That is what lets the dashboard say what to do: "reconnect this account" and "the provider is down, sit tight" are not the same message.

Who finds out

Only the user can reconnect. No amount of engineering changes that, so the design is judged by how the notice reaches them:

  • In the dashboard, on the specific account, saying what happened and what to do. A red icon with no sentence turns into a support ticket.
  • Through a signed webhook to the customer's application, when the integrator is another product. Their interface is what their user sees, and without that notice they cannot ask them for anything.
  • Silently when there is nothing to do. One alert per transient failure teaches people to ignore alerts.

What to ask before budgeting

  1. How long does each provider's permission last, and is any of them unrecoverable?
  2. Does the refresh token rotate? If it does, who holds the lock?
  3. What happens to an account nobody touches for two months?
  4. How does the user find out, and how long does it take?

None of the four is answered by reading the provider's documentation: they are answered by looking at how the system is built. It is part of what our integrations and APIs service includes, and part of why maintaining an integration costs money after delivery.

Frequently asked questions

How long does an OAuth access token last?
There is no single answer, and that is the part that breaks designs. Across the providers we maintain there are tokens lasting 60 days (Meta, LinkedIn, Threads), 24 hours (TikTok), 60 minutes (Google, and with it YouTube and Google Business), under 30 minutes (Bluesky), and permissions that never expire at all (a Discord webhook, a Slack bot token). A system that assumes one rhythm falls over on the first provider that does not follow it.
Is a scheduled job enough to refresh tokens?
Only if the token lasts much longer than the job's interval. With 60-minute tokens and an hourly pass, some accounts arrive expired at the next call, so you also need on-demand renewal right before the token is used. Both paths have to be idempotent: whichever gets there first leaves a fresh token and the other does nothing.
Why does an account disconnect itself when the token was still valid?
Almost always because of a double refresh. When the refresh token is single-use and rotates on every exchange (which is how the Bluesky protocol works), if the scheduled job and the application refresh at the same time, one of them redeems a spent token, the provider rejects it, and the account ends up flagged as broken although nobody did anything wrong. The fix is a per-account lock plus writing only if the token you started from is still the one in the database.
Where does the user come into this?
They are the only one who can fix it when it genuinely fails. A revoked or unrecoverable permission is not solved by retrying: the person has to authorise again. So the failure has to reach them with a message naming the account and the action, and that is why it pays to tell a transient failure from a permanent one before alarming anybody.