Discussion about this post

User's avatar
Josh Woodruff's avatar

The no-defenders assumption is the one I'd flag hardest. The benchmark's target is a flat network with domain admin at the end, and that's the same kind of environment the Hugging Face swarm moved sideways through in July. A segmented network with short-lived credentials changes the score without changing the model. That's the control test I'd want run next to the capability test.

No posts

Ready for more?