The discovery of 543,699 still-valid credentials across public GitHub repositories is a sobering demonstration that exposed secrets remain one of the most persistent and avoidable weaknesses in modern software development. Truffle Security scanned a snapshot of more than 224 million public GitHub repositories containing over 58 billion files and then actively validated candidate credentials against the services that issued them. The result was not a collection of strings that merely looked like passwords or API keys. The researchers say all 543,699 credentials still authenticated when tested on July 27 and 28, 2026.
That distinction is critical.
Secret-scanning tools routinely generate findings that include expired keys, test credentials, invalid tokens, examples and false positives. In this research, the headline number represents credentials that were still usable.
The median exposed credential had been sitting publicly for 784 days, meaning the typical live secret had remained available for more than two years. Ten percent were at least 6.3 years old, and the oldest valid credential had been associated with a file last modified in 2009.
That is difficult to explain as a temporary developer mistake.
A credential accidentally committed this morning is an error.
A credential remaining public and valid for years is a credential-management failure.
The research corpus came from The Stack v3, a dataset assembled from public GitHub repositories for training large language models. Its crawl closed on August 7, 2025. Truffle Security later tested candidate credentials in July 2026, meaning the researchers were validating secrets against a snapshot that was already almost a year old.
That makes the result even more striking.
These were not credentials caught moments after publication.
Many had survived publicly for years and still worked long after the repository snapshot was created.
The 543,699 unique credentials appeared across more than 1.1 million files and repository copies, including forks. This demonstrates one of the fundamental difficulties of secret exposure in source control: once a credential enters Git, removing it from one visible file does not necessarily remove every copy.
Forks may preserve it.
Commit history may preserve it.
Caches may preserve it.
Dataset mirrors may preserve it.
Search indexes may preserve it.
AI-training corpora may preserve it.
An attacker may already have copied it.
That is why the correct response to a publicly exposed credential is not:
“Delete it from GitHub.”
It is:
“Revoke or rotate it immediately, then remove the exposed copies.”
Deleting a secret from a repository addresses visibility.
Rotation addresses compromise.
Confusing those two steps leaves organizations with repositories that look cleaner while the attacker may still possess a perfectly valid credential.
The research also provides an unusually clear measure of how long this problem persists after GitHub introduced increasingly strong secret-protection features.
GitHub made secret-scanning alerts free for public repositories in 2023 and enabled Push Protection by default for public repositories in February 2024. Push Protection attempts to detect credentials before they are committed and blocks the push when a supported secret pattern is identified.
Yet Truffle Security found that 199,843 of the still-valid credentials, approximately 36.8% of the total, were exposed after Push Protection became the default.
At first glance, that might appear to suggest Push Protection is ineffective.
The actual research shows something more nuanced.
For secret types covered by the protection mechanism, Truffle Security found that the rate of leakage fell by roughly 53% after Push Protection became enabled by default.
So the feature works.
The problem is coverage.
Approximately 51.8% of the live credentials identified by researchers belonged to categories GitHub’s default Push Protection did not recognize or block, including some database connection strings and Google API credentials.
That distinction matters because it exposes a recurring problem in security tooling.
A control can be effective at everything it detects and still leave a significant blind spot.
Organizations often deploy secret scanning and then mentally convert:
“We scan for secrets”
into:
“Secrets cannot leak.”
Those statements are very different.
Detection depends on recognizable formats, provider integration, entropy patterns, validation logic and known credential structures.
A conventional AWS access key follows a recognizable format.
A custom database password embedded inside a connection string may not.
A proprietary API token may not.
A secret generated internally may look like an ordinary random string.
The challenge becomes even greater when credentials appear inside JSON, YAML, Terraform files, CI/CD configuration, notebooks, environment files, Dockerfiles, shell scripts, test fixtures or documentation.
That is why secret management should not rely entirely on post-generation pattern recognition.
The stronger approach is to avoid putting long-lived credentials into source code at all.
Applications should retrieve secrets dynamically from dedicated secret-management systems such as cloud-native vaults or enterprise secret stores.
CI/CD pipelines should receive ephemeral credentials at runtime.
Developer workstations should use short-lived access wherever possible.
Service identities should rely on workload identity or federated authentication instead of static credentials.
The best secret-scanning alert is still the one that never needs to exist because the credential was never hardcoded.
The research also exposes major differences between credential providers in how they respond after secrets leak.
BleepingComputer cites Truffle Security’s finding that of 101,886 npm tokens identified in the corpus, only one remained valid, suggesting extremely effective revocation or lifecycle controls.
By contrast, researchers found 69,041 still-valid credentials among 126,963 exposed Google Cloud service-account credentials.
That contrast is remarkable.
Both credential classes may appear accidentally in public code.
One ecosystem appears to ensure that most exposed credentials stop working.
The other leaves tens of thousands usable.
This shows that secret security is not solely the repository owner’s responsibility.
Credential providers can significantly reduce the damage from exposure by automatically detecting public leaks, revoking tokens, shortening credential lifetimes and making long-lived secrets difficult to create.
GitHub operates a Secret Scanning Partner Program that can notify participating providers when their credential patterns are discovered publicly. But provider participation does not automatically mean the credential will be revoked.
GitHub recommends treating detected secrets as compromised, but the provider ultimately controls what happens next.
Truffle Security’s research suggests that notification without revocation may leave credentials exposed for years.
This raises a useful security-design question:
If a provider knows with high confidence that one of its credentials is public, why should that credential continue to work?
There are operational reasons automatic revocation can be difficult.
Revoking a production credential may break an application.
The provider may not know whether the credential is genuinely sensitive.
Some customers may intentionally publish restricted API keys.
But the alternative is allowing a known exposed credential to remain valid indefinitely.
Where automatic revocation is too disruptive, providers could at least force administrative review, restrict privileges, shorten token lifetime, or require explicit acknowledgement before continued use.
Long-lived static credentials are themselves part of the problem.
If an API token is designed to work indefinitely until manually revoked, then one accidental commit can create years of exposure.
Short-lived credentials fundamentally change that equation.
A token valid for one hour may still leak.
But by the time somebody discovers it in an archived repository months later, it is useless.
This is one of the reasons cloud identity is gradually moving toward workload identities, federated credentials and ephemeral access tokens.
The objective is not merely to store secrets more carefully.
It is to reduce how many permanent secrets exist in the first place.
The GitHub findings also matter because public repositories are actively searched by attackers.
This is not a theoretical threat model.
MITRE ATT&CK specifically documents threat actors searching public code repositories for leaked credentials, including groups such as HAFNIUM and LAPSUS$.
Automated secret harvesting makes the attack extremely inexpensive.
A criminal does not need to compromise the developer.
They do not need phishing.
They do not need malware.
They do not need a software vulnerability.
They simply search information that the organization accidentally published.
The attack chain can be brutally short:
developer commits credential → repository becomes public → automated scanner finds key → attacker validates credential → attacker accesses cloud, database or API
In some cases, the interval between publication and discovery by attackers may be measured in minutes.
Organizations should therefore assume that a secret exposed publicly has already been copied even if the repository was corrected quickly.
The familiar argument:
“It was only public for ten minutes”
provides very little reassurance when automated bots can scan GitHub continuously.
This is why revocation needs to happen first.
Cleanup comes second.
The consequences depend heavily on the type and privilege of the credential.
An exposed low-privilege analytics key may have limited value.
An exposed cloud administrative key may provide an attacker with the ability to create infrastructure, retrieve secrets, access storage, launch compute resources or establish new identities.
A database connection string may provide direct access to customer information.
A GitHub personal access token may enable repository modification or supply-chain compromise.
A CI/CD credential may provide access to build pipelines.
A service-account credential may bypass the interactive authentication controls that protect human users.
The risk therefore cannot be measured simply by counting credentials.
Security teams need to understand the blast radius behind each credential.
That includes:
what service it authenticates to,
which privileges it carries,
whether it can create additional identities,
what data it can reach,
whether it can modify production systems,
and whether its activity is logged.
This is particularly important for machine identities.
Human accounts increasingly have MFA, conditional access, anomaly detection and session controls.
API keys and service credentials often do not.
Possession of the key is the authentication.
There is no second factor.
No user prompt.
No suspicious-login warning.
No human noticing that the account is suddenly active at 3 a.m.
If the credential works, the attacker is trusted.
That makes machine identities extremely attractive.
The research also has implications for AI development.
The dataset itself was assembled to train large language models, illustrating how exposed secrets can propagate beyond their original repositories into secondary datasets and mirrors.
Truffle Security previously found 221,303 valid credentials in Hugging Face-hosted AI training data, and the GitHub-derived corpus produced more than twice that number.
This does not mean an AI model automatically memorizes and reproduces every credential contained in its training dataset.
But it does demonstrate that public secrets can travel far beyond the repository where they originated.
Organizations therefore cannot assume that rewriting Git history completely erases exposure.
Once public data has been mirrored into datasets, archives, forks and crawls, control over distribution has already been lost.
Another important lesson is that secret scanning should extend beyond GitHub.
GitGuardian’s 2026 research estimates that roughly 28% of secret-sprawl incidents originate exclusively outside code repositories, including collaboration and productivity platforms.
Credentials leak in Slack.
Tickets.
Wikis.
SharePoint.
Email.
ChatGPT conversations.
CI logs.
Support portals.
Container images.
MCP configuration.
Documentation.
The development environment has become a large distributed information system.
Scanning only Git repositories therefore addresses one important surface, not the entire problem.
Modern secret-management programs need discovery across code, collaboration systems, CI/CD, cloud configuration and endpoints.
The incident-response process also deserves improvement.
GitHub’s own security guidance says exposed credentials should trigger investigation of what the credential actually did, including reviewing audit events, unexpected actors and unknown IP addresses.
That is critical because rotation tells you the credential no longer works.
It does not tell you whether someone already used it.
The proper response to confirmed public exposure should therefore look like:
revoke credential → issue replacement → search repository and history → determine exposure period → review provider logs → identify suspicious use → assess accessed resources → remove unauthorized changes → prevent recurrence
The audit step is frequently forgotten.
Teams rotate the key and close the ticket.
If the credential had been public for two years, that leaves a rather large unanswered question.
Organizations should also implement automated rotation wherever possible.
A credential that never changes creates unlimited opportunity for historic exposure to remain exploitable.
Regular rotation does not eliminate leaks, but it reduces the useful lifetime of forgotten credentials.
Even better is event-driven rotation.
If secret scanning detects a credential in a public repository, the system should automatically revoke or quarantine the credential and issue a replacement through a controlled workflow.
Human approval can follow.
The attacker should not receive the same waiting period.
Pre-commit controls are another useful layer.
By the time a secret reaches GitHub, the organization may already be in incident-response mode.
Local pre-commit scanning and IDE integrations can detect credentials before they leave the developer workstation.
CI checks provide a second layer.
GitHub Push Protection adds another.
Continuous repository scanning adds another.
Provider-side revocation adds another.
No individual control will catch everything, which is precisely why layered protection matters.
Organizations should also pay attention to secret-scanner bypasses.
GitHub allows developers to bypass Push Protection in certain circumstances, such as when a detected value is a false positive or intentionally published.
That capability is operationally necessary.
It can also become the corporate equivalent of clicking “Ignore” on a fire alarm.
Security teams should monitor bypass events and review them, especially for repositories belonging to production services.
Another useful metric is credential age.
A secret created six years ago and still active deserves more scrutiny than one generated yesterday.
Old credentials often belong to forgotten integrations, departed employees, abandoned projects or infrastructure nobody remembers owning.
Those credentials can become dangerous precisely because nobody expects them to be used.
Regular inventory should therefore identify long-lived keys and challenge whether they still need to exist.
Secrets should have owners.
Expiry dates.
Purpose.
Scope.
Rotation policy.
If nobody knows why a credential exists, that is not a valid reason to preserve it indefinitely.
The broader lesson from the 543,699 GitHub credentials is therefore not simply that developers make mistakes.
Everyone already knows that.
The more troubling finding is that the surrounding systems frequently fail to make those mistakes temporary.
Push Protection misses some credential types.
Providers do not always revoke exposed tokens.
Organizations fail to rotate secrets.
Repositories remain public.
Old credentials remain active for years.
Those failures compound.
The security objective should not be to build a world where nobody ever accidentally commits a secret.
That is optimistic to the point of fantasy.
The realistic objective is to build systems where:
accidental secret exposure becomes short-lived and low-impact.
A developer will eventually paste the wrong key.
A script will eventually contain a password.
A .env file will eventually be committed.
The critical question is what happens next.
Does the push get blocked?
Does the provider receive an alert?
Does the key automatically expire?
Does the SOC see unauthorized use?
Does the credential have limited privilege?
Or does it quietly remain valid until somebody discovers it 784 days later?
The Truffle Security research suggests that for hundreds of thousands of credentials, the answer was the last one.
That is the part organizations should find most uncomfortable.
The real vulnerability is not that the secret appeared on GitHub.
It is that years later, the secret still opened the door.
More than 543,000 credentials exposed in public GitHub repositories were still valid in July despite the platform's security measures to prevent accidental leaks of sensitive data. [...]
Source: Over 543,000 valid credentials exposed in public GitHub repositories via Bleeping Computer — published 30 Sep 2026.
Was this article helpful?
Your feedback helps us improve the knowledge base.