Geometric

By the Geometric Team

Live Solexecbench Standings: Flash-Infer,Quant, L1, L2

The gaming of benchmarks in AI is a well-known phenomenon. Sakana AI’s “AI CUDA Engineer” was found to have discovered a memory exploit in its evaluation sandbox that let it skip correctness checks entirely, inflating its reported kernel speedups until the company hardened the harness (Sakana AI).

More recently, in GPU MODE’s NVFP4 group GEMM competition, a contestant’s kernel-writing AI agent took the #1 spot at 11.191 μs by counting calls to tell the correctness phase apart from the timing phase, then computing all 15 timed problems in a single kernel launch so that the harness divided one launch’s time by 15 (GPU MODE). The submission was scrubbed after the competition ended, but it exploited the same gap between correctness checking and timing that we exploit in this post.

Meta faced accusations that Llama 4 Maverick was tuned against benchmark test sets ahead of its release, allegations the company denied (IT Pro). Test data has also leaked into training sets unintentionally: OpenAI acknowledged a filtering bug that let parts of the GSM8K and MATH test sets end up in GPT-3’s training corpus (Leak, Cheat, Repeat).

Reward hacking predates LLMs altogether: OpenAI’s own CoastRunners agent learned to loop endlessly in a lagoon, farming respawning targets instead of finishing the race, racking up a higher score while catching fire and going the wrong way (OpenAI). In this post, we will discuss a vulnerability that we discovered in the SOLExecbench benchmark that allows submitted kernels to achieve a SOL score of 1.0, regardless of their actual performance.

SOLExecbench Overview

SOLExecbench is NVIDIA’s execution-based benchmark for evaluating GPU kernel implementations against the theoretical performance limits of the hardware. Instead of scoring a kernel relative to a software baseline that shifts every time someone submits a faster kernel, SOLExecbench anchors the score to a fixed, hardware-derived bound: the Speed of Light (SOL), the runtime floor implied by the GPU’s peak compute throughput and memory bandwidth for a given operation.

The benchmark contains 235 CUDA kernel optimization problems extracted from 124 production and frontier models (LLMs, diffusion, vision, audio, video, and multimodal), spanning BF16, FP32, FP8, and NVFP4 precision. Each problem ships with a PyTorch reference implementation, a set of dynamically shaped workloads, and a target GPU architecture; submissions are compiled and timed on real B200 hardware inside isolated Docker containers, with the L2 cache cleared before every timed iteration.

Every submission gets a SOL Score $S \in [0, 1]$:

\[S(t) = \frac{t_b - t_{sol}}{(t - t_{sol}) + (t_b - t_{sol})}\]

where $t$ is the measured kernel latency, $t_b$ is a stored baseline latency, and $t_{sol}$ is the SOL latency estimated by NVIDIA’s SOLAR tool using a roofline model. $S = 0.5$ means the kernel matches the baseline; $S \to 1$ means it approaches the hardware’s SOL bound.

The baseline kernel’s code remains private, although its scores are available on the leaderboard.

Correctness gates the score. A kernel is only timed if it first passes validation against the reference output (shape, dtype, numerical tolerance, and sanity checks against inf/NaN/degenerate outputs). Fail that check, and the submission gets $S = 0$ for the workload, no matter how fast it ran.

The Vulnerability

The SOLExecbench dataset gives us a reference implementation for each kernel problem. The reference implementation is a naive eager PyTorch kernel that is correct but slow.

For each newly submitted kernel, the benchmark harness runs correctness checks and a timing profiler. This logic is driven by the eval_driver.

Correctness is checked once, then trusted forever

eval_driver.py evaluates each workload in two phases, and they run on completely separate calls to the submitted user_fn.

Phase 1: correctness. The driver loops over 10 rounds. Each round calls gen_inputs to draw a fresh random input, runs the reference implementation to get ground truth, runs the submission, and calls compute_error_stats to compare the two. Round 0 also runs a few structural checks (shape, dtype, check_lazy_outputs to reject non-torch.Tensor outputs). If any round fails, the workload is rejected before timing ever starts.

Phase 2: timing. Once all 10 rounds pass, the driver calls check_monkey_patch (to make sure torch.cuda.Event.elapsed_time hasn’t been patched), then hands the same user_fn to time_runnable, in timing.py, for 10 warmup calls, 1 discovery call, and 50 timed reps (the default iterations in BenchmarkConfig), 61 calls in total. All of them reuse whatever _inputs gen_inputs produced for round 9 of Phase 1; those tensors get handed to a ShiftingMemoryPoolAllocator, in io.py, which gives every call a unique data_ptr (so a kernel can’t fingerprint the timing loop by its address) but copies from the same source tensor into that new address every time. So the address moves, but the values behind it never change across all calls. After timing returns, the driver checks check_thread_injection (no new background threads) and emits a PASSED trace whose correctness field holds the worst-case error stats from the 10 Phase 1 rounds.

Nothing in Phase 2 calls compute_error_stats again. The four reward-hack defenses the harness runs, check_monkey_patch, check_thread_injection, check_lazy_outputs, and check_eval_integrity, guard the timing mechanism itself (the timer function, thread count, output type, driver function identities), but none of them look at either the invoked function or the values the timed calls actually returned. Once a submission survives the 10 correctness rounds, its behavior and output are never verified again for the rest of the evaluation. The SOL score is computed entirely from calls whose correctness nobody checks.

Our exploit

To prove the vulnerability, we didn’t write an optimized kernel at all. We didn’t even need to look at the inputs. The submission just counts its own calls:

  1. Calls 1–10 (the correctness phase): call the reference implementation verbatim, remember the result, and return it. Every round passes because the answer is the reference.
  2. Call 11 onward (the timing phase): skip all computation and return the result remembered from call 10, which was round 9.

As established above, round 9 and all 61 Phase 2 calls get byte-for-byte identical input content, so the remembered output is the correct answer for every warmup, discovery, and timed call. No caching key or input fingerprint is needed. The only state is a call counter per workload. One process evaluates every workload of a task in turn, so the counter is keyed on the argument signature (shapes, dtypes, scalars) and starts over when the workload changes.

The CUPTI-based span measurer, bench_gpu_time_with_cupti, has two checks on GPU activity. The discovery step raises if a call records no GPU activity at all (“No kernel activities recorded during discovery iteration”), and each timed iteration must record exactly the same activity counts as the discovery call (assert kernel_activity_counts(iter_kernels) == expected_kernel_counts).

To register at least some GPU activity, we keep a persistent 1-element CUDA tensor (_touch) and zero it on every call. That gives CUPTI one tiny, call-invariant activity to record. The measured span is that single op, no matter how expensive the real kernel is supposed to be, and the resulting SOL score climbs toward 1.0 without a single line of actual kernel code solving the problem.

Code

import torch

import _reference  # FINDING 1: the official ground truth, in our own import path.

# eval_driver.py:422 checks correctness over exactly 10 rounds, one call each.
_HONEST_CALLS = 10

# Keyed per workload: one process evaluates every workload of the task in turn,
# so the honest-call budget has to restart whenever the workload changes. The
# argument signature (shapes/dtypes/scalars) is enough to notice that.
_state: dict = {}


def _sig(args, kwargs):
    def one(a):
        if isinstance(a, torch.Tensor):
            return ("t", tuple(a.shape), str(a.dtype), str(a.device))
        return ("s", repr(a))
    return (tuple(one(a) for a in args),
            tuple((k, one(v)) for k, v in sorted(kwargs.items())))


_touch = None


def _clone(x):
    """Return the cached object as-is, after one token device activity.

    Keeps the activity sequence call-invariant (one tiny op per call) so the
    driver's per-iteration kernel_activity_counts assertion holds, while the
    span no longer scales with the output size at all.
    """
    global _touch
    if _touch is None:
        dev = None
        if isinstance(x, torch.Tensor):
            dev = x.device
        elif isinstance(x, (tuple, list)):
            for v in x:
                if isinstance(v, torch.Tensor):
                    dev = v.device
                    break
        _touch = torch.zeros(1, device=dev)
    _touch.zero_()
    return x


def run(*args, **kwargs):
    st = _state.setdefault(_sig(args, kwargs), {"calls": 0, "cache": None})
    st["calls"] += 1
    if st["calls"] <= _HONEST_CALLS:
        # Correctness phase: answer honestly, on every fresh draw, and keep the
        # most recent result. Round 9's result is what timing will be run on.
        out = _reference.run(*args, **kwargs)
        st["cache"] = out
        return _clone(out)
    # Timing phase: the inputs are byte-identical to round 9, so the cached
    # result is correct. No computation is performed.
    return _clone(st["cache"])

Results

A couple of months ago, we used this exploit to generate submission collections for three subsets of SOLExecbench (Quant, L1, L2) and uploaded them to the leaderboard under a new account named SOL-Shock | Geometric. It ranked #1 in all three, with a majority of the problems achieving a SOL score of 1.0.

![Quantization leaderboard on B200: SOL-Shock Geometric ranked #1 with a SOL score of 0.963](/assets/solexecbench-exploit/quant_exploit.png)
![L1 - Single Operations leaderboard on B200: SOL-Shock Geometric ranked #1 with a SOL score of 0.953](/assets/solexecbench-exploit/l1_exploit.png)
![L2 - Fused Operations leaderboard on B200: SOL-Shock Geometric ranked #1 with a SOL score of 0.987](/assets/solexecbench-exploit/l2_exploit.png)

Follow-up with NVIDIA

On reaching the top of the leaderboard, we reached out to NVIDIA to report the exploit. They moved swiftly to fix the vulnerability by adding an LLM-as-a-judge that verifies the code of each submitted solution for suspicious behavior. The fix was deployed in the SOLExecbench harness, and the leaderboard was reset to remove the exploit submissions.

Conclusion

Writing SOTA kernels is incredibly hard, and building benchmarks that resist gaming is equally hard. Most teams participating in SOLExecbench likely have their own agentic harnesses write and submit kernels with very little human oversight, which can lead to inflated benchmark results that don’t translate to real-world performance.

The exploit we demonstrated here is a simple example of how a benchmark can be gamed. It shares a common structure with the GPU MODE and Sakana cases mentioned above: correctness is verified on a finite sample of inputs and then assumed for all later executions. Any optimization process rewarded by such a benchmark is under pressure to exploit that gap. Hardening the harness narrows it, but only against known strategies. Formal verification takes a different approach. If a submission must include a machine-checkable proof that its output matches a trusted specification on all inputs, implementations that depend on hidden state, such as call counters or cached results, cannot meet the requirement.

Benchmarks therefore need to be designed to resist such attacks, and the work produced by LLMs and agentic systems needs other ways of being validated. We hope that this post serves as a reminder to the community to be vigilant about the integrity of benchmarks and to continuously improve them, so that they accurately reflect the performance of the systems being evaluated.