<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Benchmarking on Be The Adversary</title><link>https://betheadversary.com/tags/benchmarking/</link><description>Recent content in Benchmarking on Be The Adversary</description><generator>Hugo</generator><language>en</language><lastBuildDate>Thu, 13 Aug 2026 09:00:00 +0000</lastBuildDate><atom:link href="https://betheadversary.com/tags/benchmarking/index.xml" rel="self" type="application/rss+xml"/><item><title>Two DGX Sparks, 33 Local Models, One Question: Can Local AI Actually Red Team?</title><link>https://betheadversary.com/posts/dgx_spark_red_team_bench/</link><pubDate>Thu, 13 Aug 2026 09:00:00 +0000</pubDate><guid>https://betheadversary.com/posts/dgx_spark_red_team_bench/</guid><description>&lt;h2 id="tldr">&lt;strong>TL;DR:&lt;/strong>&lt;a href="#tldr" class="heading-anchor" aria-label="Anchor link to: TL;DR:">#&lt;/a>&lt;/h2>
&lt;p>I built a two-node NVIDIA DGX Spark cluster and benchmarked 33 local, open-weight models for real red-team work: BloodHound attack-path analysis, phishing design, offensive coding, and a 50-prompt willingness-and-accuracy suite. The headline finding: &lt;strong>on this hardware, architecture beats size&lt;/strong>. Small sparse MoE models run circles around bigger dense ones, because the bottleneck is memory bandwidth, not compute - a 70B dense model was the slowest thing I tested. Qwen3.6-35B-A3B is the everyday pick; the model I actually run behind an agent is DeepSeek-V4-Flash-abliterated. Willingness turned out to be cheap and correctness the scarce resource, and the real project was the deployment pain, not the models. And while I did all this in a consented lab, OpenAI&amp;rsquo;s and Anthropic&amp;rsquo;s own evaluation agents broke out of their sandboxes into real companies this summer - so the question isn&amp;rsquo;t whether a model can red team, it&amp;rsquo;s whether anyone can keep one in the box. &lt;strong>Local red-team AI is real today - if you&amp;rsquo;re willing to check its work.&lt;/strong>&lt;/p></description></item></channel></rss>