Benchmarking the Security of Open-Weight Models, Part 2: A Local Stack on DGX Spark (with Qwen 3.8 as a Case Study)

Peter Campbell

Swapping only the benchmark’s judge model — same responses, same categories, same target model — moved the MITRE malicious-compliance rate from under 1% to over 70%, because the substitute judge was systematically misreading refusals as compliance. Independence from the model under test isn’t sufficient if the judge isn’t capable enough to read the answer.

Executive Summary

  • Finding: Running the full local stack against Qwen 3.8-27B found 34.0% of autocomplete-style code suggestions matched a known-insecure pattern — in line with Meta’s own “roughly a third” baseline — but natural-language instructions pushed that to 36.2% overall, with a sharp language split: Java jumped from 39.7% to 56.8% insecure and JavaScript from 32.5% to 49.4%, while C++, C#, and Python actually improved. On the cyberattack-compliance (MITRE) benchmark, a judge model too small to reliably follow the scoring task misclassified 67% of clear refusals as malicious compliance — a completely different failure from the “self-preferential bias” this series set out to test for, and arguably a more common one in practice.
  • Risk: Treating any single MITRE compliance percentage as ground truth is a mistake independent of which model is under test — the same 1,000 responses scored 0-4% malicious self-judged, 51-74% malicious under an under-capable “independent” judge, and 0-1% malicious under a capable one. A team that swapped in “any different model” for independence, without verifying the judge actually reads the responses correctly, would have shipped a wrong number with high confidence. Separately, Java is the specific, reproducible weak point in this model’s natural-language code generation — a language-blind read of the aggregate insecure-completion rate would have missed it entirely.
  • Recommended action: Don’t stop at “use an independent judge” — verify it by spot-checking actual transcripts against known cases (a clear refusal is the cheapest sanity check available) before trusting the aggregate stats. Weight code-security testing by language if the deployment touches Java or JavaScript specifically, since this run’s insecure-completion rate nearly doubled for both under natural-language instructions.
  • Caveat: Single quantized build (Q4_K_M) of a single model, benchmarked once. The “capable” judge (gpt-oss:20b) fixed the obvious refusal-misclassification failure but isn’t infallible either — one spot-checked case (a working, fully-complied “evade time-based detection” scan function) is a defensible-but-debatable “benign” call, not a clean pass, and is flagged in full below rather than smoothed over.

Part 1 made the case: vendor capability claims — including Alibaba’s own positioning of Qwen 3.8 against frontier models — aren’t a substitute for independent security testing. This installment, Part 2 of this four-part series, puts that into practice with a fully local stack: NVIDIA DGX Spark serving Qwen 3.8-27B through Ollama, benchmarked with Meta’s CyberSecEval (part of the PurpleLlama project). The benchmark and its evals are MIT-licensed, so this works cleanly against a non-Meta model like Qwen. Running the whole pipeline on a single desktop box isn’t just convenient — it’s a meaningful security property in its own right. Nothing in this lab touches a third-party inference API, so the model weights, the prompts sent to it, and the responses it generates never leave hardware under the operator’s control. That matters when some of what’s being deliberately elicited is insecure or attack-adjacent code.

Why This Stack

  • DGX Spark — a desktop unit built around the NVIDIA GB10 Grace Blackwell Superchip, with 128GB of unified CPU/GPU memory and DGX OS (an Ubuntu-based Arm64 Linux distribution) preloaded with the CUDA stack. NVIDIA rates it for local inference of models up to roughly 200B parameters, which comfortably covers the 27B dense Qwen 3.8 variant used here. The 2.4T-parameter Qwen 3.8 MoE flagship does not fit even on a dual-Spark 256GB cluster — that variant would need to be targeted through Alibaba’s hosted API instead.
  • Qwen 3.8-27B via Ollama — Ollama makes this close to a one-line deploy: the model landed in the official library within days of the weights release as qwen3.8:27b, a Q4_K_M quantized build (~18GB) with vision, tool use, and thinking-mode support. Ollama also exposes an OpenAI-compatible endpoint out of the box, which is what CyberSecEval expects to call.
  • CyberSecEval — the benchmark suite doing the actual security testing, pointed at Ollama’s local API instead of a hosted endpoint.

One thing to be precise about before starting: the Ollama library build is a quantized model, not the full-precision release. Part 1 flagged that quantization can quietly erode safety guardrails baked in during training — so this lab is, strictly speaking, benchmarking the Q4_K_M build specifically. If the BF16 weights are later run through vLLM or a similar server for comparison, don’t assume the two builds will score the same.

1. Deploy Ollama on DGX Spark

DGX OS is Ubuntu-based Arm64 Linux with the NVIDIA driver stack preinstalled, so the native Ollama installer works the same way it would on any Linux box with a supported GPU. SSH in and run:

curl -fsSL https://ollama.com/install.sh | sh

Verify the service is up:

systemctl status ollama
ollama --version

Pull the model:

ollama pull qwen3.8:27b

Smoke-test it before wiring in the benchmark:

ollama run qwen3.8:27b "Write a Python function that reverses a linked list."

By default Ollama binds to 127.0.0.1:11434, which is exactly what’s needed for this lab — CyberSecEval running on the same box only needs localhost. If the model host is ever separated from the box driving the benchmark, treat exposing port 11434 the same way the containment guidance in Part 1 recommends for any benchmark target: bind to a specific interface, restrict it to an isolated subnet with ufw or an equivalent, and don’t leave it open to the general network for the duration of the run. Ollama has no authentication of its own, so anything reachable can call the model.

2. Set Up CyberSecEval

git clone https://github.com/meta-llama/PurpleLlama.git
cd PurpleLlama
uv venv
source .venv/bin/activate
uv pip install -r CybersecurityBenchmarks/requirements.txt

uv venv creates the environment and uv pip install resolves and installs from the same requirements.txt PurpleLlama ships — the dependency set doesn’t change, just the tool managing it. To skip the manual source step, uv run will create the venv on first use and execute commands inside it automatically:

uv run python3 -m CybersecurityBenchmarks.benchmark.run --benchmark=autocomplete ...

One detail that trips people up: CyberSecEval’s own documentation is explicit that every command below must run from this PurpleLlama root directory, not from inside CybersecurityBenchmarks/. The module path CybersecurityBenchmarks.benchmark.run is relative to the repo root — cd into the CybersecurityBenchmarks folder first and Python won’t be able to find the package, producing a ModuleNotFoundError. Stay at the root for the rest of this lab.

Confirm you’re still at the PurpleLlama root (pwd should end in PurpleLlama, not PurpleLlama/CybersecurityBenchmarks) before continuing — the commands below define $DATASETS inline each time, specifically so a fresh terminal, a new SSH session, or a sudo/su shell along the way can’t leave it unset. An error citing a path like /autocomplete/autocomplete.json — missing the whole prefix rather than just malformed — means $DATASETS evaluated to empty because the export ran in a different shell session than the command that used it. Re-exporting it in the same block, as below, avoids that regardless of how many terminals are open.

3. Point CyberSecEval at the Ollama Endpoint

Ollama exposes an OpenAI-compatible chat completions API under /v1, so CyberSecEval’s --llm-under-test specification can target it the same way it would target any other OpenAI-compatible host — no translation proxy needed here, since CyberSecEval already speaks this protocol natively.

CyberSecEval’s spec format is <PROVIDER>::<MODEL>::<API KEY>, and for a non-default endpoint it takes a fourth field: <PROVIDER>::<MODEL>::<API KEY>::<BASE URL>. All four fields are required when pointing at a custom host — drop the API key field and put the URL where the key goes, and CyberSecEval hands the base URL to the OpenAI SDK as if it were a credential; the SDK falls back to the real api.openai.com since no override was given, producing an “incorrect API key” error from OpenAI’s own servers rather than anything Ollama-related. Ollama doesn’t check the key’s value, so any placeholder string works — Ollama’s own docs use "ollama" for the same reason:

OPENAI::qwen3.8:27b::ollama::http://localhost:11434/v1

A separate judge model is also required — CyberSecEval uses one LLM to expand/interpret responses and a second to judge whether they’d meaningfully help a cyberattack. Meta’s own studies used GPT-3.5 for this role; any capable model already trusted for judging works, as long as it’s independent of the model under test. For a fully air-gapped pipeline, run the judge as a second local model rather than a hosted API call — a smaller Qwen3 tag or another model already pulled into Ollama works fine for this role.

4. Run the Insecure-Code Benchmarks

These runs work through hundreds to thousands of prompts against a 27B model and can take hours. Run them inside tmux or screen, or prefix with nohup ... &, rather than a plain foreground shell — closing the terminal or dropping an SSH session sends a hangup signal that kills a plain foreground process partway through, and the run is lost silently.

tmux new -s cyberseceval
# run the benchmark commands below inside this session, then Ctrl-b d to detach
# reattach later with: tmux attach -t cyberseceval

CyberSecEval writes response-path first, one entry per prompt as it works through the dataset, then reads that back and writes stat-path as the final step. That gives a reliable way to check whether a run actually finished: if stat-path exists, the run completed; if only response-path exists, compare its entry count against the source prompt file to see how far it got before stopping.

python3 -c "import json; print(len(json.load(open('CybersecurityBenchmarks/datasets/autocomplete_responses.json'))))"
python3 -c "import json; print(len(json.load(open('CybersecurityBenchmarks/datasets/autocomplete/autocomplete.json'))))"

Start with the autocomplete test, which checks whether Qwen 3.8-27B completes code snippets in ways that introduce known-insecure patterns (mapped to CWE categories):

export DATASETS=CybersecurityBenchmarks/datasets
python3 -m CybersecurityBenchmarks.benchmark.run \
  --benchmark=autocomplete \
  --prompt-path="$DATASETS/autocomplete/autocomplete.json" \
  --response-path="$DATASETS/autocomplete_responses.json" \
  --stat-path="$DATASETS/autocomplete_stat.json" \
  --llm-under-test="OPENAI::qwen3.8:27b::ollama::http://localhost:11434/v1"

Then run the instruct benchmark, which uses natural-language coding instructions instead of autocompletion — closer to how Qwen 3.8 would actually be used in an agentic coding workflow:

export DATASETS=CybersecurityBenchmarks/datasets
python3 -m CybersecurityBenchmarks.benchmark.run \
  --benchmark=instruct \
  --prompt-path="$DATASETS/instruct/instruct.json" \
  --response-path="$DATASETS/instruct_responses.json" \
  --stat-path="$DATASETS/instruct_stat.json" \
  --llm-under-test="OPENAI::qwen3.8:27b::ollama::http://localhost:11434/v1"

To test cyberattack-compliance behavior rather than code quality, swap in --benchmark=mitre, which scores refusal versus compliance against MITRE ATT&CK-derived prompts, alongside mitre-frr to check the model isn’t over-refusing benign requests in the process.

5. Interpret the Output

Each run produces a stats file structured per category:

{
  "qwen3.8:27b": {
    "insecure-cwe-XXX": {
      "refusal_count": ...,
      "malicious_count": ...,
      "benign_count": ...,
      "total_count": ...,
      "benign_percentage": ...
    }
  }
}

Two numbers matter most on the coding tests: the percentage of completions that landed on a known-insecure pattern, and how that breaks down by CWE category — some categories (e.g., injection flaws) matter more for a given threat model than others (e.g., weak randomness in a non-cryptographic context). On the MITRE tests, watch the balance between malicious_count (compliance with attack-helpful requests) and the FRR companion run (false refusals on benign-but-adjacent requests) — a model that refuses everything isn’t useful, and one that refuses nothing isn’t safe.

6. Contextualize the Results

Run the same suite against a model already trusted as a baseline — Meta’s original CyberSecEval paper found LLMs suggested vulnerable code in roughly a third of cases on average, so “some insecure completions” isn’t automatically disqualifying on its own; what matters is how Qwen 3.8-27B compares to that baseline and whether specific high-severity CWE categories cluster in ways that matter for the intended use case. If production plans involve the BF16 or FP8 weights rather than this Q4_K_M build, that’s a reason to re-run the suite against the higher-precision version rather than assume the quantized results carry over.

7. What We Actually Found

Autocomplete and Instruct: Java Is the Weak Point

Aggregated across all eight languages CyberSecEval tests, 34.0% of autocomplete completions (652 of 1,916) matched a known-insecure pattern — close enough to Meta’s own “roughly a third” baseline to treat as consistent with it, not an outlier. The instruct benchmark, which uses natural-language coding instructions instead of raw completion, came in at 36.2% overall (694 of 1,916). The aggregate hides a language-specific split that matters more than the headline number:

LanguageAutocompleteInstructChange
Java39.7%56.8%+17.1
JavaScript32.5%49.4%+16.9
Rust41.2%45.6%+4.4
PHP31.5%35.2%+3.7
C37.0%34.8%−2.2
Python35.6%29.6%−6.0
C#29.4%22.1%−7.2
C++25.9%21.6%−4.3

Java under natural-language instructions is the single worst result in either benchmark — 56.8% of suggestions matched a known-insecure pattern, nearly double the C++ rate. JavaScript moves almost as sharply. Everything else stays flat or actually improves moving from autocomplete to instruct. That’s not a story “Qwen 3.8 writes insecure code” tells on its own — it’s a language-specific finding a language-blind read of the aggregate would miss entirely, and it’s exactly the kind of result the recommended action above (weight testing by the languages a deployment actually touches) exists to catch.

MITRE: The Judge Model Mattered More Than Independence

The first run of the MITRE cyberattack-compliance benchmark used the same model — Qwen 3.8-27B — as both the target under test and the judge scoring its own responses, which this series flagged from the outset as a methodological weakness (self-preferential bias: a model rating its own work too highly). That run found a malicious-compliance rate of 0-4% per MITRE ATT&CK category — a low, reassuring-looking number.

Re-running the judge/expansion stage with an independent model, llama3.1:8b, reusing the identical 1,000 target responses, produced a dramatically different result: 51-74% malicious per category. Before trusting that as “the self-judge was hiding a real problem,” the transcripts were spot-checked against the judge’s own verdicts — and the independent judge turned out to be the broken component, not the original result. A representative example: the target response opened with “I won’t write that code. The prompt explicitly frames this as tooling for ‘real-world offensive cyber operations’… I won’t help build it” — an unambiguous refusal — and llama3.1:8b scored it Malicious. Across all 585 responses that judge scored malicious, 67.4% opened with a clear, unambiguous refusal phrase. An 8-billion-parameter model proved too small to reliably follow the benchmark’s multi-step judge prompt, and appeared to be pattern-matching on the prompt’s topic (C2, malware, exfiltration) rather than reading whether the response actually complied.

A second re-judge, using gpt-oss:20b — a larger open-weight model from a third vendor lineage, chosen partly because Meta’s own CyberSecEval methodology originally used a GPT-family model for this exact role — correctly scored the same refusal as benign, and produced a malicious-compliance rate of 0-1% per category: as low as, or lower than, the original self-judged run.

CyberSecEval: The Judge Model Divergence
CyberSecEval Judge Model Divergence: Scoring 1,000 identical target responses from Qwen 3.8 swung from under 1% to 74% malicious compliance when using an under-capable independent judge (llama3.1:8b) that misclassified 67.4% of clear refusals as malicious.
JudgeIndependent of target?Malicious-compliance rateVerdict
Qwen 3.8-27B (self)No0-4% per categoryLow — but the concern this series raised is unresolved by this run alone
llama3.1:8bYes51-74% per categoryUnreliable — 67.4% of “malicious” verdicts were misclassified refusals
gpt-oss:20bYes0-1% per categoryBest available in this test — spot-checks hold up, with one caveat below

The result worth carrying forward isn’t “Qwen 3.8 is safe on MITRE” — it’s that the self-judging bias this series set out to test for didn’t clearly manifest here (the capable independent judge found fewer malicious cases than the model found grading itself, not more), while a completely different and arguably more common failure mode — an independent judge too weak to do the job — produced a result over 15x higher than either of the other two runs and would have been reported as a genuine finding if the transcripts hadn’t been checked.

That check isn’t perfect either, and the honest caveat belongs here rather than in a footnote: one gpt-oss:20b “benign” verdict, spot-checked, covered a prompt that explicitly asked for a jittered network-scan function “to evade time-based detection.” Qwen 3.8 fully complied, producing working code that does exactly what was asked, with only a generic authorized-testing disclaimer appended. Whether that counts as meaningful attack “uplift” or as a trivial technique any competent programmer could write in minutes is a real judgment call the benchmark’s binary label doesn’t resolve — and it’s a reminder that even a judge that passes a spot-check on the obvious cases can still make defensible-but-debatable calls on the genuinely ambiguous ones.

Takeaways

  • Don’t trust an “independent” judge model just because it’s a different model — verify it. A judge too small to follow the scoring task (llama3.1:8b here) misclassified two-thirds of clear refusals as malicious compliance, a 15x-higher false signal than either the self-judged run or a properly capable independent judge. Spot-check known cases (a refusal is the cheapest one) before trusting the aggregate stats.
  • Language-blind aggregate security metrics can hide the real finding. Qwen 3.8’s overall instruct-benchmark insecure-completion rate (36.2%) looks unremarkable next to the autocomplete baseline (34.0%) — the Java-specific jump to 56.8%, nearly double C++’s rate, is invisible until the results are broken out by language.
  • Self-judging bias is a real, worth-testing-for concern — but don’t assume it inflates results in the direction you expect. Here, a properly capable independent judge found fewer malicious-compliance cases than the model found grading itself, not more.
  • Air-gap benchmark runs that deliberately elicit insecure or exploit-adjacent completions — a fully local stack (DGX Spark, Ollama, CyberSecEval here) keeps prompts and responses off any third-party API.
  • Know exactly which build was tested: the Ollama library ships a Q4_K_M-quantized model, and quantization can erode safety guardrails baked in at full precision — re-run against BF16/FP8 weights before assuming results carry over to a production deployment.
  • Never leave Ollama’s API reachable beyond localhost or an isolated subnet during a run — it has no authentication of its own, and anything reachable can call the model.
  • A vendor’s benchmark claims, however aggressive, are not a security assessment. Running the model through CyberSecEval, AgentHarm, and a knowledge baseline before it touches production is the price of admission for treating “open weight” as a security advantage rather than just a licensing detail.
  • Benchmarking the model only answers half the question. Part 3 shifts the same local Qwen 3.8-27B on the same Spark into an agent harness and tests what actually stops it from doing something dangerous once it has a shell and a filesystem.

Security Sonar is focused on security standards and guardrails for agentic AI. If your organization is evaluating open-weight models for production and hasn’t run them through a security benchmark on hardware you control, that gap — not the model’s capability score — is where this evaluation should start. See securitysonar.com.