Fix scan failures
A scan that fails to finish leaves your Retrievy Index running on the last successful set of findings, so the dashboard ages instead of breaking. This page walks you through what to do when you see a Scan Failed badge on a data source or get the Scan Failed email.
Before you start
- A role with the right permission for the source you need to fix: Create, edit, and delete cloud accounts (cloud sources), Create, rotate, and revoke agent tokens (agent-driven sources), or Manage FortiGate device connections (FortiGate devices). Each lives under its matching group in Settings → Roles & Permissions. Without the right toggle, you see the data source but the action buttons are disabled.
- The failure email Retrievy sent to your team, or access to the affected data source in Settings.
- For agent-relayed scans (Active Directory, FortiGate via agent), the host running the Retrievy Agent online. Check Fleet → Agents first if you're not sure.
Quick triage
When a scan fails you'll see two signals.
The first is an email with the subject [Retrievy] Scan Failed [account name] ([scan type]). The red header reads Scan Failure, followed by the workspace name, the data source that failed, the failure timestamp in your workspace timezone, an Error category in red caps, and the raw technical message. The View Configuration button at the bottom of the email drops you on Fleet → Agents as a starting point.
The second is a red Scan Failed badge on the data source's card or row inside Retrievy:
- Cloud accounts: Settings → Cloud Credentials, on each provider card.
- FortiGate devices: Settings → FortiGate Firewalls, on each device row.
- Windows / Active Directory: Settings → Windows Servers, on each domain row.
Open the failed card. Triage in this order:
- Read the Error category in the email to narrow down where to look (credentials, network, configuration).
- If the source is agent-relayed, check Fleet → Agents for the agent's status before touching credentials. An Offline agent will fail every scan it's asked to run.
- Click Test Connection on the source. Retrievy runs a fast credential and reachability probe and surfaces a green or red message inline. Fix what the probe complains about, then run a new scan.
Reading the failure message
The email's Error category maps to a known failure shape. The most common ones, and what they mean:
| Error category | What happened | Where to look first |
|---|---|---|
| Authentication / Authorisation Error | The API token, key, or service principal was rejected by the data source. | Settings card for the source. Re-enter the credential, then Test Connection. |
| Network Timeout | Retrievy or the agent couldn't reach the device on the required port within the timeout. | Firewall, NAT, security group, or routing between the agent and the device. |
| Connection Error | The host or endpoint is unreachable: refused, no route, name resolution failed. | IP address or hostname, management interface state, recent network changes. |
| SSL / TLS Error | The TLS handshake failed: expired certificate, version mismatch, self-signed without the right toggle. | Device certificate, Verify SSL toggle on the source card. |
| Configuration or File Error | A scan started but the agent couldn't read the config file or backup, or the upload didn't arrive. | Agent host read access, Run Scan retry from Settings. |
| Incomplete Scan Output | The scanner finished but returned no usable findings. | Often a transient issue. Retry once before digging further. |
| Partial Scan Failure | The scan finished for some checks but a subset failed mid-run. | The technical message names the failing module. Retry, then escalate if it keeps recurring. |
| Unexpected Error | The scanner threw something that doesn't match the categories above. | Open the Technical Details code block in the email, then contact Retrievy support if it isn't obvious. |
Two more failure messages come from Retrievy itself, not the data source:
- "Scan timed out while pending in queue." The scan sat in the Pending state for more than 30 minutes without progress. Usually means the workspace was at the concurrent scan cap (see Scan rules) and the job aged out. Trigger a new scan once active work clears.
- "Scan exceeded maximum execution time (4h)." The scan ran for more than 4 hours and was force-failed. Means the scanner stalled, not that your data source is huge. Retry once. If it recurs on the same source, contact Retrievy support with the data source name.
The cleanup pass that produces these two messages runs every 5 minutes, so a stuck scan is failed and the badge updates between 30 and 35 minutes after it stalled (or between 4 hours and 4 hours 5 minutes for the long timeout).
Cloud account failures
Cloud accounts (AWS, Azure, GCP, OCI, Microsoft 365, Cloudflare) fail in a small number of repeatable ways. Walk down this list before deleting the account.
The credential rotated. Someone in the cloud provider's console rotated the access key, secret, client secret, API token, or service-account JSON. Retrievy is still holding the old value.
Fix: open Settings → Cloud Credentials, click the affected card, and re-enter the new credential in the field that changed. Click Test Connection. When the green confirmation appears, click Scan Now.
The IAM policy was revoked or narrowed. The cloud-side identity Retrievy uses lost a permission a scanner check needs. The failure category will be Authentication / Authorisation Error, but the Test Connection result is usually green because the probe only needs a tiny subset of the read permissions a full scan needs.
Fix: re-attach the full read-only policy for the provider. The exact IAM policy documents live on each integration page: AWS, Azure, GCP, OCI, Microsoft 365, Cloudflare. After reattaching, click Scan Now.
A Retrievy cloud account expects a non-interactive identity. If someone enables MFA on the IAM user, service principal, or service account Retrievy is using, every scan will fail with Authentication / Authorisation Error. Either remove MFA from that identity or create a dedicated non-interactive identity for Retrievy. Never share a human user's MFA-protected credentials with the platform.
The region or subscription went down or got renamed. Less common but real: a regional outage on the provider's side, or someone moved the resource to a different subscription / project.
Fix: confirm the provider's status page is green. Check that the region or subscription ID on the source card still matches reality. Run Scan Now once the provider clears.
The OAuth consent was revoked. Specific to Azure / Microsoft 365 Quick Connect. An admin in the tenant clicked Revoke admin consent on the Retrievy app registration.
Fix: re-run the Quick Connect flow from Microsoft 365 or the Azure setup page. Retrievy creates a fresh consent grant. The old data source row is replaced when the new one finishes onboarding.
Agent-driven failures
Active Directory and FortiGate-via-agent scans run on the Retrievy Agent. A failure on these almost always lives on the agent host, not on Retrievy's side.
The agent is offline. The agent has missed more than 20 minutes of heartbeats. Every scheduled scan that targets this agent will fail until it reconnects.
Fix: open Fleet → Agents and look at the agent row. If the badge is Offline or the agent shows Not responding on the detail panel, work through agent troubleshooting first. The cloud account or FortiGate device that depends on this agent will start succeeding again on its next scheduled scan once the agent comes back.
The agent's outbound network broke. The agent host can reach the local data source fine but can't reach the Retrievy workspace, or the proxy credentials on the host changed.
Fix: on the agent host, confirm outbound HTTPS to your workspace domain works. If you set a proxy at install time, re-check the proxy credentials in the agent's environment. The agent system tray (Windows) shows the connection state if you need to read it locally.
FortiGate is behind a NAT or the IP changed. The configured host or port for the device no longer routes from the agent.
Fix: open Settings → FortiGate Firewalls, click the three dots on the device row, then Edit Details. Update the host, port, and any NAT-specific settings. Save, then trigger Force Scan Now from the same menu.
DNS or hostname resolution is broken on the agent host. The agent host can't resolve the data source's hostname.
Fix: on the agent host, test name resolution against the configured hostname. If you're using an internal DNS, confirm the agent host points at it. If the device only has an IP, set the FortiGate or AD source to use the IP directly.
SSH or LDAP credentials were rotated. Specific to FortiGate (SSH) and Active Directory (the read-only service account). The local credential rotated but the agent still has the old one.
Fix: open the source in Settings, click Edit Details, re-enter the credentials, save. The agent picks up the new credentials on its next heartbeat. Then trigger a scan.
Force Scan Now and Run Scan both queue a command on Retrievy that the agent picks up on its next heartbeat. If the agent is Offline, the command sits in the queue and runs whenever the agent comes back. Use Fleet → Agents to confirm the agent's state before chasing a failed scan retry.
When the scanner stops responding
A small subset of failures show up with the message "Capsule lost (no status, no output) after [N]s deadline." This means the scan was started but never reported a result. Retrievy waits up to 90 minutes before giving up and marking the scan Failed.
You don't fix this on the data source. Trigger a new scan from Settings → Cloud Credentials with Scan Now. The same message twice in a row on the same source is worth contacting Retrievy support, with the data source name and the failure timestamps in hand.
Manual retry
Once you've fixed the underlying cause, run a new scan instead of waiting for the next scheduled cycle.
Cloud accounts
Open Settings → Cloud Credentials. On the failed card, click Scan Now. The badge flips to a blue Scanning... pill with a spinner. When the scan finishes, the badge becomes a green Scan Success.
Settings → Cloud Credentials → [account card] → Scan Now
FortiGate firewalls
Open Settings → FortiGate Firewalls. On the device row, click the three dots on the right, then Force Scan Now.
Settings → FortiGate Firewalls → [row] → ⋯ → Force Scan Now
Active Directory
Open Settings → Windows Servers. On the active domain card, click Run Scan. The button switches to Scanning... while the agent picks up the job on its next heartbeat.
Settings → Windows Servers → [domain card] → Run Scan
A specific agent
Open Fleet → Agents, click the agent row to open its detail panel, then click the lime Run Scan at the top of the panel. A toast confirms Scan queued for [hostname]. Starts on next heartbeat. The agent picks the job up within ten minutes.
When the next scheduled scan fires
If you'd rather wait, Retrievy queues one scan per eligible data source every day at 23:00 in your workspace timezone. A workspace whose subscription is in a paused state (past due, unpaid, paused, canceled in grace) does not run scheduled scans. The full table of which states scan and which don't lives at Scan rules → Plan eligibility.
Last resort: re-add the source
Deleting and re-adding a data source is rarely necessary. Almost every failure resolves by fixing the credential, the network path, or the agent. Re-adding has a real cost: you lose the data source's scan history, drift timeline, and any Security Exceptions still scoped to it.
Re-add only when one of these is true:
- The cloud subscription or tenant changed and the new identity will never reconcile against the old data source row (for example, the AWS account moved to a different organisation).
- The data source was added with the wrong configuration mode (manual paste when Quick Connect was the right choice) and you want to start clean.
- Retrievy support tells you to.
To re-add: open Settings, locate the source, click Delete, confirm. Then run the relevant onboarding wizard again from the same page. Findings from the previous source are removed; the new source starts a fresh scan history once its first scan completes.
A delete wipes every finding, drift event, and exception scoped to the source. The action is permanent. Try the credential and network fixes above before deleting.