Research

DOM-0.8B

Model card for DOM-0.8B, a local code-security model for severity and vulnerability-family scoring.

🤗 Hugging Face | 📑 Blog Post | 💡 GitHub

Introduction

Today we’re introducing DOM-0.8B, a 760M parameter code-security model, alongside vakt, our local scanner. Our mission is to make security review accessible to everyone, regardless of nationality or budget. It’s free, open weights and fully local, under the Vaktex Evaluation License (see LICENSE): free for research and education, with commercial use available on request. You can now optimize your code review without using cloud models.

Our tool vakt extracts functions from your repository, scores them with DOM-0.8B and ranks them for further investigation. Rather than generating a written review of every function, DOM returns severity and vulnerability-family scores. The aim is to focus human review and generative-model tokens where they are most useful, before the code reaches production.

Getting Started

Prerequisites:

  • macOS on Apple Silicon, or Linux (amd64/arm64, glibc 2.35+)
  • A Hugging Face account with access to vaktex/dom-oss-0.8b (it’s gated: request access on the model page)

Install

curl -fsSL https://get.vaktex.com/oss-vakt | sh

Download Weights

hf auth login
vakt summon

Scan the current directory

vakt .

Find more functionality on our GitHub

Architecture

DOM-0.8B uses a 24-layer text backbone with a hidden width of 1,024. Each of six groups combines three linear-attention layers with one full-attention layer. In the current recipe, the lower 16 layers remain frozen while the upper eight are adapted.

DOM-0.8B architecture diagram

One backbone pass per input chunk produces token representations. Learned four-head attention pooling in fp32 combines them into a representation for a severity score and 18 CWE family scores. A separate 256-dimensional training projection is discarded for inference.

A generative model processes the input and then decodes its answer one token at a time. DOM processes each input chunk and returns classification scores, without generating a written report. That removes answer-generation work and gives vakt fixed outputs it can rank and threshold. The engineering advantage is a narrower task; its effect on review time and accuracy still needs to be measured.

Benchmarks

Severity AUROC measures how well a model ranks vulnerable code above non-vulnerable code across thresholds. Higher is better; it is not percentage accuracy.

DOM-0.8B vs DOM-4B benchmarks

An AUROC of 0.5 represents chance-level ranking; 1.0 represents perfect separation on the evaluated examples. DOM-4B’s score of 0.8371 means a randomly chosen vulnerable example ranks above a randomly chosen non-vulnerable example about 84% of the time, counting ties as half. It does not mean 84% of findings are correct.

DOM-0.8B scores 0.5731, below the character n-gram baseline at 0.6009. That baseline uses recurring text patterns rather than explicit program analysis, making it a useful check on what the model adds. The graph reports a stronger result for DOM-4B, but establishing an improvement requires evaluating the builds on the same held-out examples. For review workflows, precision and recall at the chosen threshold also matter: they determine how much code gets flagged and how many issues are missed.

Best Practices

  • Route on families. The family probabilities are the strongest signal. Reading the top one or two is a reasonable triage convention, and a snippet can legitimately belong to several.
  • Set thresholds on your own code. The right cut-off depends on how you weigh a missed vulnerability against a false alarm. Choose it from a labelled sample of your own repositories, per family where volume allows.
  • Use it to prioritise, not to decide. A high probability is a reason to look, not a finding. Treat the output as an ordered queue for human or downstream review.
  • Score one unit at a time. The model sees one function or file. It cannot reason about a flaw whose cause sits in a caller, a configuration file or another service.
  • Split very large files. Most files fit within 16,384 tokens, but predictions above roughly 8,192 tokens are less well supported. Scoring by function gives steadier results.
  • Always set the language tag. Coverage is strongest in mainstream server-side languages and weaker elsewhere, including smart contracts.

License

DOM-OSS and the vakt tool are released under the Vaktex Evaluation License (see the LICENSE file). The weights are open and source-available: anyone may download, inspect, run, and study them.

  • Research and education are free: including publishing benchmarks and papers, teaching, and personal learning.
  • Personal or internal evaluation is free: to decide whether to license it commercially.
  • Commercial or production use requires permission from Vaktex. Companies and individuals can request it by requesting access on the model card. Downloading the weights does not, by itself, grant the right to deploy them or use them in the course of a business.
  • Derivatives (fine-tunes, quantizations, distillations) are permitted for research and education, provided they carry this same license.
  • No resale and no public API hosting, ever. Even with commercial permission, you may not resell the model or tool, nor host it behind a public API or as-a-service offering to third parties. Internal-only use is fine. These limits change only under a separate written agreement signed by Vaktex.

For terms beyond these, or access to Vakttårn for deeper review, contact [email protected].

Improve team velocity with
better security and privacy.