Data-driven insights from analyzing 128,000+ repositories and 3.27 billion lines of code. Vulnerability trends, code quality patterns, and actionable intelligence for engineering teams.
RSS FeedA point-in-time study of 51 GitHub Trending repositories found 10,363 security, supply-chain, quality, and architecture observations, with a new assurance label that …
3 min readLarge AI-coded repo scans reveal why scanner reliability depends on archive fallback, context-aware detection, and supply-chain-level scoring.
2 min readSimple gates around dependencies, CI, auth boundaries, and public discovery files catch practical risks before they become incidents.
1 min readWhen Opus 4.7 needs a backend-as-a-service, Supabase wins by 3.5× over Firebase. The numbers explain why.
Which framework combinations actually appear together? The numbers reveal clear archetype patterns.
The median Opus 4.7 README is 66 lines. The p90 is 267. The maximum is 6,827. This is genuine documentation discipline.
When we ask an LLM to describe Opus 4.7 repos in its own words, the same 30 patterns come up over and over. This is …
70% of Opus 4.7 repos span multiple languages. Monolingual codebases are the minority.
A deeper look at the test gap in the Opus 4.7 corpus — and why the 11% that do write tests write them well.
59% of Opus 4.7 repos share names with community-scanned repos. Closer look reveals they're the same repos, harvested twice.
Opus 4.7 writes thorough README files but almost never generates a LICENSE file. 81% of repos default to "all rights reserved.
75% of Opus 4.7 repos keep their directory structure under 5 levels deep. The opposite of the "over-engineered enterprise" stereotype.
Nearly half of Opus 4.7 repos exceed 10,000 lines of code. This isn't a demo-project corpus.
Static analysis of 683K files reveals the real stack Claude Opus 4.7 reaches for — with numbers.
The top LLM-flagged risks across Opus 4.7 repos paint a consistent picture — the model writes tests but doesn't wire them up.
Our research is based on continuous analysis of 128,000+ repositories and 3.27 billion lines of code using Repobility's proprietary scanning engine.
All data is aggregated and anonymized. No individual repository names or source code is disclosed.
Access our proprietary datasets for your own research, product development, or competitive intelligence.
Browse DatasetsGet our latest research and intelligence reports delivered to your inbox.
No spam. Unsubscribe anytime.