EVM Spend Challenge
Under EIP-8288, a privacy pool could prove its spend rules with a zkVM proof of a program's execution. The program could be EVM bytecode, executed by an EVM interpreter inside the zkVM, or a RISC-V program compiled from Rust. We wrote the same spend check both ways and proved both with leanVM. How much more the EVM program costs to prove depends mostly on how its bytecode is executed, and on whether leanVM has an instruction for the hash precompile that the EVM program calls.
What was measured
The spend check follows the minimal shielded pool's rules: two input notes in a 20-level Merkle tree, two output notes, a public amount and a fee. It makes 55 hash calls, 105 compressions of 64-byte blocks in total. The RISC-V program computes each hash either with base instructions, the standard RV64IM instructions that correspond to EVM opcodes, at about 3,100 cycles per BLAKE2s compression, or with leanVM's blake2s instruction, which computes one compression as a single instruction and so corresponds to an EVM precompile. The EVM program calls a hash precompile, which the interpreter computes with base instructions or with the blake2s instruction. Without a precompile, the EVM program computes the hash with EVM opcodes.
Results
Select a result to see the measurements behind it in the explorer below.
Each ratio divides a program's value by the baseline's, drawn as a dashed line; click any row to make it the baseline. Cycles count the instructions leanVM executes, the largest count over the measured valid test cases; a blake2s instruction counts as one cycle but costs more to prove, which proof times reflect. Proof times are medians of four to six proofs of one spend on an Apple M5 Max, without zero knowledge. The two hash editions were timed in separate runs; retained samples of the same program differ by up to about 1.9×, so small timing differences are inconclusive. Gas is function execution gas under this restricted Cancun profile, excluding transaction overhead; the hypothetical BLAKE2s precompile is priced like SHA-256. RISC-V programs have no gas metric here.
What this means for EIP-8288
With a dedicated hash instruction, the interpreter for proving takes about twice the RISC-V cycles on this spend. With BLAKE2s on both sides, it takes 1.96 times the direct RISC-V cycles and 2.31 times the median proof time; REVM takes 3.85 times the cycles. What counts as close enough is still open. These measurements cover a restricted, stateless spend function, without zero knowledge. They do not establish the cost of arbitrary EVM applications or production private proofs.
A shared interpreter could bind many programs under one verification key. It would take bytecode as input and bind its hash, execution profile and output. This benchmark instead embeds each program's bytecode. The current EIP-8288 draft identifies generic STARK relations by verification-key hash; a contract that pins one hash cannot automatically accept another. A proposed execution scheme could instead bind the bytecode hash and execution profile, with protocol upgrades replacing the interpreter and verifier while preserving those semantics. That scheme and its upgrade rules are additional design work. The interpreter becomes part of what verifiers trust and needs corresponding review and testing.
The fast hash path depends on the exact primitive exposed. L1 has SHA-256 at 0x02 and KECCAK256. The pinned leanVM has a BLAKE2s instruction; L1's BLAKE2 precompile at 0x09 computes BLAKE2b compression, a different function. EVM programs run under L1's EVM rules, so a BLAKE2s or BLAKE3 call would require a new L1 precompile. Any pool that uses such a hash, whether its proved program is EVM bytecode or RISC-V, also needs it on L1, or a separate proof, to update its tree and to recompute the statement digest it matches against the proof's public-input hash. SHA-256 already matches L1. A Keccak-f instruction could accelerate KECCAK256 with Ethereum's padding; a complete SHA3-256 instruction is not automatically compatible, because SHA3-256 and Keccak-256 use different domain separation.
Other heavy primitives need their own measurements. The measured BLAKE2s bytecode takes 17 to 223 times the cycles of the base-instruction RISC-V version and about 26,000 gas per compression, averaged over this spend. A second version, which keeps four 32-bit words in each 256-bit word, takes about 11,700 gas per compression and about half the interpreted cycles, but its executions still exceed leanVM's single-proof limit. These results measure two implementations of 32-bit arithmetic in EVM bytecode. They do not establish that every signature or hash is practical only with an L1 precompile. Changing the hash also changes commitments and nullifiers, so the SHA-256 and BLAKE2s editions are different relations, not interchangeable implementations of one deployed pool.
RISC-V remains the long-term target. Direct RISC-V avoids EVM interpretation and can implement new primitives with base instructions. EIP-8288's generic interface does not require one universal RISC-V key. Shared verifier parameters with program-hash binding already exist in other zkVMs, such as RISC Zero; compatibility with the proposed aggregation scheme is a separate question. This benchmark does not implement that integration or measure a stable program-hash interface across prover upgrades.
The challenge
The challenge asks for the EVM bytecode that checks this spend in the fewest cycles, with SHA-256 as the statement hash and 0x02 as the only allowed precompile. Everything else is fixed: REVM executes the bytecode, and the leanVM version and SHA-256 code do not change. The score is the largest cycle count over the valid test cases, including new ones generated at scoring time. Three entries so far:
| Cycles | vs RISC-V | Gas | |
|---|---|---|---|
| RISC-V reference (Rust) | 460,471 | 1.00× | |
| Solidity baseline | 1,225,610 | 2.66× | 41,075 |
| Optimized Yul | 625,721 | 1.36× | 14,658 |
| Hand-written bytecode | 615,971 | 1.34× | 14,559 |
Plain Solidity pays for code the compiler generates around each hash call. The two hand-optimized entries come close to a program that only makes the 55 precompile calls, at 1.22 times the RISC-V cycles, so the rest of the difference is the interpreter's cost. Gas ranks the entries in the same order but overstates their differences: the Solidity baseline uses 2.82 times the gas of the hand-written bytecode but only 1.99 times its cycles. The SHA-256 precompile's gas covers about 85 proving cycles per unit, against about 18 for the other opcodes under REVM.
The specification gives the rules, input format and scoring. The challenge has not launched.
How it was measured
The prover is leanVM's riscv-exploration branch at commit 1096dedf: RV64IM plus the blake2s instruction, without zero knowledge. The harness executes a stateless EVM function with only the hash precompile, no storage, logs or environment reads, a 30-million gas limit and a 1 MiB memory limit. Native fixture tests check valid digests and invalid-input rejection. The fast interpreter was also compared with REVM on about 3.6 million random programs; the compiled entries were checked on fixed fixture sets at multiple gas limits. These are finite tests, not a proof of equivalence or a production security audit. Each guest embeds its bytecode. Hashing this bytecode as dynamic input is estimated to add about 10,000 cycles with the blake2s instruction; the complete dynamic-input path has not been measured. Results and raw measurements document the checks and their scope.