Potential impact
What could happen if the underlying issue were exploited: loss of funds, key exposure, code execution, privacy loss, or availability impact.
Three separate questions: How risky might the change be? How clearly does its public message explain it? And how well does the vendor practice security in public?
The score expresses the security relevance of one public change—not the overall safety of a project. Every point comes from a visible dimension.
What could happen if the underlying issue were exploited: loss of funds, key exposure, code execution, privacy loss, or availability impact.
Required access, complexity, user interaction, preconditions, and whether exploitation is practical at scale.
Whether abuse could remain invisible and whether the commit language obscures or clearly states security relevance.
How widely the changed component is deployed and whether vulnerable state persists after a software update.
How strongly the available commit, diff, tests, issue links, and surrounding context support the classification.
The directness and completeness of primary-source evidence.
Immediate, credible risk of catastrophic harm.
Serious impact or practical exploitation likely.
Meaningful security relevance with limiting conditions.
Defense-in-depth or narrow, hard-to-exploit exposure.
Security-adjacent change with no credible active risk found.
The deterministic 0–100 score evaluates the public commit message—not the code quality, developer skill, or safety of the patch. It is calculated for every captured commit and keeps its reasons and warnings alongside the score.
Every message starts at 20 before positive evidence and penalties are applied.
Up to 30 points for a descriptive, concrete subject and 12 for a recognizable type or component scope.
Context beyond the title earns up to 23 points. Rationale or a described failure mode adds 12 more.
Testing or verification earns 10, issue/advisory references earn 8, and explicit security behavior earns 5.
Generic placeholders can lose 45 points, very short subjects 20, two-or-fewer-word subjects 10, and work-in-progress language another 20. The result is clamped from 0 to 100.
Specific purpose with meaningful supporting context.
The change is identifiable, though some context may be absent.
Some purpose is visible, but rationale or evidence is limited.
The public record does not adequately explain the change.
A message such as “runs” can be poor public evidence even when the patch is harmless. CommitWatch flags the communication gap without adding points to the security-risk score.
Five equally weighted dimensions—disclosure quality, researcher acknowledgement, security process, patch clarity, and incident response—are each scored from 0 to 100 and averaged. Missing evidence is not automatically failure; it is labeled insufficient data until a defensible score exists.
A strong product can have weak disclosure practices. A transparent vendor can still ship a flaw. CommitWatch keeps product risk, message clarity, and vendor behavior separate.
Fetch the public author string, message, diff, file list, metadata, and verified references. Deterministic triage and message-quality scoring run on every commit.
The configured Ollama model produces structured risk dimensions and plain/technical summaries for ranked candidates. Successful output publishes immediately with unavoidable AI provenance.
Researchers, vendors, and readers can submit community notes that add evidence, qualify a claim, or identify an error.
A human moderator validates notes before they appear. Approved context remains visibly separate from raw AI analysis, and material corrections update the public record.
Git author names are public strings and may not uniquely identify a person. Message quality measures documentation, not developer competence. A diff is evidence, not omniscience: private reports, unreleased patches, hardware behavior, operational controls, and vendor context may change the conclusion.
Corrections & challenges →