Tool credential health
Credential expiry reminders, OAuth token refresh and rotation, reconnect prompts, and why a server looks offline
Overview
A tool connection stops working for one of three reasons: the credential expired, the OAuth grant was invalidated, or the server itself is not reachable. MagOneAI tracks each one separately and tells you which it is.
This page is about keeping connections alive. For setting one up, see External MCP servers.
Static credential expiry
A bearer token, API key or service-account secret expires on a date only the issuing system knows. MagOneAI cannot discover it, so you declare it.
When you connect a template that takes a sensitive field and is not OAuth, you can set a credential expiry date. MagOneAI then reminds you before it passes.
| When | What happens |
|---|---|
| 7 days before | A reminder |
| 1 day before | A reminder |
| After the date passes | A notice that it has expired |
Recipients are the person who created the connection, plus organization owners and admins for project-scoped and organization-scoped connections. A departed creator therefore does not mean a silent expiry.
An expiry date never takes a connection offline. The date is something you typed and MagOneAI has not verified, so a mistyped year must not break a working connection. Expiry informs, it does not block.
The reverse is also true: a credential that expires without a date set gives you no warning at all. The connection simply starts failing.
Changing the date re-arms every reminder, because each reminder is recorded against the expiry it referred to. Correcting a mistyped date gives you the full warning sequence again.
If the reminder email does not arrive
Notification mail always sends from your organization's own email credentials, never from a platform address. An organization with no email provider configured gets the in-app notice and the badge on the connection row, but no email.
That is deliberate: falling back would put platform-address mail in front of your users without your organization opting in. See Notifications.
Rotating a credential
Rotating rewrites the secret at the connection's existing storage path, so the connection ID survives and every workflow that references it keeps working. You do not rebuild the connection.
Rotation also clears the expiry reminder record, so set the new date at the same time.
Rotation is refused for OAuth connections. Re-run consent instead, which is the OAuth equivalent.
OAuth tokens
The platform owns OAuth. It runs consent, holds the refresh token, and knows when the access token dies, so an MCP server receives a token that already works rather than the materials to mint one.
Refresh is automatic
Every path that calls a tool refreshes the token if it is due: an agent, a workflow Tool node, and Test Connection.
The trigger is the stored expiry of the access token, which is stamped when consent completes and re-stamped on every refresh.
Refresh token rotation
Some providers issue a new refresh token on every refresh and invalidate the one just spent. Two callers refreshing at the same moment would therefore both spend the same token, leaving a dead one stored and forcing you to re-consent.
MagOneAI takes a lock per connection. The caller that waits re-reads the stored credential and uses the winner's token instead of spending its own.
Servers can write a rotated token back
A server that refreshed a rotating token can post the replacement back to MagOneAI. The secret is stored, and the new expiry lands in the column that every refresh decision reads.
This is what keeps the platform's view of the token correct when the server, rather than the platform, did the refresh.
offline_access is never narrowed away. It is what makes a refresh token be issued at all, and it is not an app permission, so it does not appear on a provider's permissions page.
An administrator filling in granted scopes from that page correctly omits it. MagOneAI exempts it from scope narrowing, because without that exemption consent succeeds, the connection works for one access-token lifetime, and re-consent is the only way back.
Reconnect prompts
A connection can show needs reauth. It means the OAuth grant is no longer valid, and it comes from two independent signals.
When consent completes, MagOneAI records which OAuth app issued the token. If the app that resolves today is a different one, every token the old app issued is orphaned.
This is worked out per request rather than stored, so there is no set of rows to keep in sync and the answer cannot go stale.
A mismatch is reported only when the record exists, the current app resolved successfully, and the two genuinely differ. Every other combination reports nothing, because a credential-store outage lighting up every connection on every tenant would be worse than the problem it reports.
An administrator who regenerates a client secret in the provider's own console changes nothing MagOneAI stores, so the check above stays silent forever. The only evidence is the refresh being rejected.
MagOneAI records that rejection, but only for a genuine dead-grant error from the provider. A 5xx, a timeout or an unreadable response is the provider having a bad day, not a dead grant.
The flag clears as soon as something proves the grant works: completing consent, a successful refresh, or a successful Test Connection.
Test Connection refreshes the token first rather than reporting on a token that is about to die. Otherwise it could send you away reassured with a connection that stops working within the hour.
Changing an organization's OAuth app
When you add or remove your organization's own OAuth app credentials, MagOneAI tells you how many connections now need reconnecting, counted after the change.
Removing them is not a neutral act. Resolution falls back to the global app, which orphans every token your organization's own app issued.
The warning before the change is generic, and the count afterwards is exact. A precise number beforehand would mean posting candidate credentials to a preview endpoint for a number you see one click later anyway.
PKCE
PKCE is off by default and enabled per MCP server, for providers that require it. Salesforce External Client Apps, for example, reject a plain authorization-code exchange.
When it is on, MagOneAI generates a verifier and challenge pair at authorize time, sends only the challenge with the S256 method, and replays the verifier when exchanging the code. The verifier is a handshake secret and is stripped from anything the API returns.
Only S256 is supported. The weaker plain method is not.
Turning PKCE on for a server that already has connections needs its auth configuration unlocked first. Registration refuses a changed auth configuration while the lock is set. See Auth config locking.
Why a server looks offline
MagOneAI checks server health every minute. A server with no heartbeat for three minutes is marked unhealthy, and an unhealthy server's tools are dropped from runs.
Superadmins are notified when a server goes offline, and again when it recovers.
Vendor-hosted servers are probed, not waited for
A vendor's own MCP server is reachable over the internet and knows nothing about MagOneAI, so it will never send a heartbeat. Left alone it would go stale, its tools would disappear, and it would look permanently disconnected with a perfectly healthy server on the other end.
MagOneAI probes those instead. Two rules make the probe meaningful:
| Response | Read as |
|---|---|
| Anything below 500, including 401 | The server is there |
| 5xx | The server is not serving |
A 401 counting as alive is the important one. The probe asks "is the server there", not "are our credentials good". Conflating the two would mark a live server dead the moment a token lapsed, which is a credential problem with a completely different fix.
A 5xx is the exception, because that is usually a gateway answering on behalf of an origin that is down.
Probing applies only to genuinely vendor-hosted servers reached over the public internet. MagOneAI's own servers heartbeat themselves, and external servers reached through the connector are deliberately not probed: the connector stops heartbeating a server whose upstream rejects it, and probing the always-available connector would hide exactly that signal.
Diagnosing a broken connection
Check whether the server is healthy
An unhealthy server affects every connection to it. Nothing you change on one connection will help.
Check for a reconnect prompt
Needs reauth means the grant is dead. Re-run consent. No amount of refreshing fixes it.
Check the credential expiry badge
An expired static credential needs rotating, with a new expiry date set.
Run Test Connection
It refreshes the token first, so a pass means the connection genuinely works now, and it clears a stale reconnect flag.
Check the tool selection
A connection that works but offers fewer tools than you expect is a scope or selection question, not a credential one. See External MCP servers.
Troubleshooting
Cause: No expiry date was set on the connection. MagOneAI cannot discover it.
Fix: Set the date when you rotate the credential.
Cause: The opposite should happen. Changing the date re-arms every reminder.
Fix: Check the new date is in the future. A date already in the past gives you the expired notice rather than the countdown.
Cause: The organization has no email provider configured. Notification mail always sends from your organization's own credentials.
Fix: Configure an email provider for the organization.
Cause: Working as designed. Tokens issued by the old app are orphaned by the new one.
Fix: The response to the credential change tells you how many connections are affected. Tell those users to reconnect.
Cause: A refresh was rejected with a dead-grant error, most often because a client secret was regenerated in the provider's console.
Fix: Re-run consent. Test Connection clears the flag once the grant works.
Cause: Unlikely to be real. The mismatch check reports nothing unless the current app resolved successfully, specifically so that a credential-store outage cannot light up every tenant.
Fix: If you see it at scale, treat it as an infrastructure problem rather than reconnecting everything.
Cause: The probe got a 5xx, which is read as not serving because it usually means a gateway answering for a dead origin.
Fix: Check what the server's URL returns directly. A 401 would count as healthy, so a 5xx is a real signal.
Cause: The server is marked unhealthy, and unhealthy servers are excluded from the tools a project can use.
Fix: Fix the server's health. This is not a per-connection problem.
Cause: The server's auth configuration may be locked, which happens once it has active connections.
Fix: An administrator unlocks it. The lock is there to stop a configuration change breaking every existing connection at once.