PIP-3: FP PoUW Certificates
Specifies a new protocol based on floating-point matrices.
| Author | Pearl Team |
|---|---|
| Status | Draft |
| Type | Standards Track |
| Category | Consensus |
| Created | 2026-07-30 |
| Source | pip-0003.md |
Table of contents
Abstract
Pearl Proof-of-Useful-Work (PoUW) certificates (versions 1/2; see PIP-2 [1]) prove integer noised matrix multiplications: operands are committed as INT8 values with one bit of headroom (INT7). This PIP specifies a new protocol in which the operands are floating-point matrices. Henceforth, the previous and new protocol are referred to as the INT Protocol and the FP Protocol respectively.
Certificate version 3 becomes mandatory at the activation height: version 1 and 2 certificates are rejected from this height onward. We stress that version 3 certificates are not valid under version 1 or 2 validation rules. Activation is therefore a consensus hard fork: upgraded nodes MUST parse and verify version 3 certificates according to this specification, while non-upgraded nodes will reject blocks using this certificate version.
Motivation
The FP protocol aims to facilitate useful work.
- Floating Point Support. The overwhelming majority of AI models are natively floating-point. The FP protocol is designed to work directly with floating-point matrices, whereas the previous protocol used integer-valued matrices.
- Low Overhead. The FP protocol enables an almost perfect “Two-For-One” incurring significantly smaller overhead compared to the INT protocol.
Caveats
We highlight several caveats and differences from the existing INT protocol:
- Lossiness. Unlike the INT protocol, the output of the FP protocol does not match the expected output in a bit-exact fashion. Despite these errors at the matrix-multiplication level, empirically, the models we tested do not exhibit noticeable differences in accuracy.
- Quantization. The INT protocol replaced the single atomic operation of multiplying two matrices, whereas the new FP protocol replaces a slightly more involved operation: quantization to FP8 (E4M3) followed by matrix multiplication. We expand on the relevance of this operation in the context of AI inference below.
- Hardware and Implementation Dependence. Different architectures, and different implementations yield different outputs (bitwise). Therefore, the miner is required to declare the hardware it uses in advance, and both the miner and the verifier follow a specific computation on the declared hardware.
Quantization
Modern inference alternates between low-precision and high-precision data types:
- Matrix Multiplication: Input matrices
A, Bare in a low-precision data type. The multiplication accumulates into a high-precision accumulator (e.g.,FP8 @ FP8 -> FP32). - Quantization: Output
Cis converted back to the lower-precision representation.
In essence, this pattern appears in virtually every modern inference pipeline; we will not delve into the technical motivation behind it.
High-Level Overview
The FP protocol exploits the quantization operator in two ways:
- Security. Quantization is a non-linear map whose behavior has no exploitable structure - computing a product of quantized matrices offers no shortcut even when the quantized matrices are low-rank. This is the computational unpredictability of the hardness assumption in the section Security Assumption.
- Accuracy. Models tolerate quantization noise - in many cases by design [4-6]. The FP protocol only increases the existing quantization noise by a small factor.
We describe a simplified version of the FP protocol below.
# Q is a quantization operator, and @ denotes matrix multiplication.
# NA, NB are low-rank noise matrices (in the full protocol NA = E1 @ F1, NB = E2 @ F2).
A' = Q(A + NA) # m x k, quantized
B' = Q(B + NB) # n x k, quantized
C'' = (A' - NA) @ (B' - NB).T # m x n
The output does not equal the expected true matrix multiplication Q(A) @ Q(B).T, but rather is very close to it (as in general, Q(A) ~ A but not equal). As the noise matrices are chosen to be of constant low rank (e.g., r = 16, independent of the other parameters), the computation of C'' is not much slower than that of Q(A) @ Q(B).
Specification
Certificate Version
This PIP assigns Pearl PoUW certificate version 3 to the FP protocol.
certificate := version || bytes
version := 3
As in prior versions, certificate bytes are not part of the block identifier.
The FP Protocol
The protocol consists of a prover (miner) and a verifier. An honest prover takes operands (matrices) A, B and performs the following computation:
0. INPUT: A in BF16(m x k), B in BF16(n x k)
1. PRE-QUANTIZATION: A, B ARE CAST TO A MEDIUM PRECISION (TBD)
2. COMMITMENT:
# H is a hash function
- COMMIT B to H(B), and A to H(A)
3. NOISE GENERATION:
- GENERATE E1 in FP8(m x r) FROM H(A) and H(B)
- GENERATE F1 in FP8(r x k) FROM H(B)
- GENERATE E2 in FP8(n x r) FROM H(B)
- GENERATE F2 in FP8(r x k) FROM H(B)
4. QUANTIZE:
# @ denotes matrix multiplication of the input dtypes using FP32 accumulator
# Q is a quantization operator (row-wise) defined below to FP8
- A' = Q(A + E1 @ F1) # m x k, quantized
- B' = Q(B + E2 @ F2) # n x k, quantized
5. MATRIX MULTIPLICATION & LOTTERY:
- C' = A' @ B'.T
- CHECK FOR WINNING TICKETS ON TILES OF C'
6. PEEL:
# || denotes concatenation of matrices horizontally
- A'' = [E1 || A' @ F2.T ] # m x 2r
- B'' = [(E2 @ F2 - B') @ F1.T || -E2 ] # n x 2r
- C'' = A'' @ B''.T
7. OUTPUT C' + C''
With the added peeling correction term, the reader may verify the output is mathematically equivalent to the formula from the High-Level Overview section, showing it approximates the true matrix multiplication up to the rounding error due to the quantization.
The verifier architecture remains unchanged, adapted to verify the above computation (ZK). Note that the verifier is required to be able to reproduce floating-point computation on the specific committed hardware in a bit-exact fashion. Hardware is therefore whitelisted - this proposal supports the Hopper and Blackwell architectures. As a consequence the miner is also restricted to a specific computation (the baseline implementation; see Reference Implementation).
Commitment
As in prior versions, H(A) and H(B) are commitment hashes binding the operand bytes, their shapes, the mining configuration, and the chain state, and support openings of individual rows. For more details on how the commitment is implemented, see the Pearl whitepaper [7].
Version 3 adds the device specification to the mining configuration: the commitment binds the whitelisted hardware (Hopper/Blackwell) on which the computation MUST be reproducible.
mining_config_bytes := common_dim
|| rank
|| mma_type // 0 = INT7 x INT7 -> INT32, 1 = FP8 x FP8 -> FP32
|| rows_pattern.subtile
|| cols_pattern.subtile
|| moe.e
|| moe.top_k
|| rows_pattern.grid
|| cols_pattern.grid
|| reserved
|| device // 0 = Hopper, 1 = Blackwell
// Lottery ticket layout is now (tile, grid).
Noise Generation
All noise factors are deterministic functions of the commitments: E1 is derived from H(A), H(B); E2, F1, F2 from H(B). Draws MUST come from the canonical sampler fed with a keyed BLAKE3 [3] stream. Every implementation MUST reproduce the draws in a bit-exact fashion.
The distribution is real-valued and symmetric, normalized per row so that the noise is small yet not negligible relative to the values subjected to it: rows of E1, E2 are drawn with an expected squared norm determined by a committed noise scale, and columns of F1, F2 with expected unit squared norm.
Quantization
Row-wise quantization: each row is quantized with a single floating-point scale stored in a higher precision (BF16). The scale is a deterministic function of the corresponding row such that the scaled noised row fits the FP8 (E4M3) dynamic range (maximum 448.0 [2]). Elements are divided by the scale and rounded to the nearest representable FP8 value.
Block Resolution
The output A' @ B'.T is partitioned into tiles committed in advance by the miner (see mining config above). Each tile is compressed by a committed extractor into a short message whose keyed Blake3 [3] hash is the lottery ticket; a ticket wins if it falls below the consensus target scaled by the tile’s arithmetic work, as in prior versions. Unlike the INT protocol, we do not have a transcript - the winning tickets are computed only on the single matrix product C' (no intermediate computations are needed). Winning tickets are subject to an additional verifier-side check (the Jackpot Policy below): statistical tests on the opened tile that reject degenerate or predictable noised computations.
Jackpot Policy: TBD
Rationale
Hardware Dependence
Hardware dependence is a necessary evil in floating-point protocols. The future seems to be abundant with various floating-point data types that provide infinitely many different flavors of accuracy/speedup tradeoff, e.g., the already ubiquitous MXFP8, NVFP4. While the FP protocol is not agnostic to the data type, this proposal is a first step toward a protocol supporting further useful workloads.
Why determinism instead of tolerance checks?
Any tolerance in verification becomes vulnerable to grinding attacks. E.g., the miner can perturb the output and get many tickets without doing additional computation.
Pre-Quantization
We allow pre-quantizing the matrices in a more memory-efficient format, thus reducing the overhead (specifically, of the commitment). Exact details regarding the format will be determined before this PIP is promoted to Proposed.
Lottery
The tickets are the tiles of the intermediate C'= A' @ B' rather than that of the final output C' + C''. This gives the miner more flexibility in the peeling process.
Committed Noise
We highlight that the noise matrices F1, E2, F2 depend solely on B, and only E1 depends on A. This might look odd at first because E1, F1 are the noise factors of A and E2, F2 are that of B. The motivation is amortization over B'': in the peeling step (see “The FP Protocol” section), B'', which is the second peeling part, depends on all three noise factors F1, E2, F2, and B'. In the context of useful work, B stands for the weight matrix and is fixed throughout a block. Therefore, this design choice enables the computation of B'' per block (rather than per matrix multiplication).
Backwards Compatibility
This PIP is not backwards compatible with nodes that only understand Pearl PoUW certificate versions 1 and 2. Blocks using certificate version 3 require upgraded consensus validation.
Security Assumption
For simplicity, we focus on the case where the adversary provides the zero matrices. In this case, the protocol reduces to computing
Q(E1 @ F1) @ Q(E2 @ F2).T
The underlying security assumption is that computing the matrix multiplication of quantized low-rank matrices is as hard as for general (unstructured) matrices of the same dimensions.
Reference Implementation
A baseline miner implementation (quantization, noise generation, matrix multiplication) will be provided so that users can test their own implementations against it, and will be linked here before this PIP is promoted to Proposed.
Test Vectors
TBD.
Experiments
TBD. We will provide experimental results on the reference miner before this PIP is promoted to Proposed.
Deployment
TBD.
References
- PIP-2: Grouped-GEMM Proof-of-Useful-Work for Mixture-of-Experts
- FP8 Formats for Deep Learning - defines the FP8 (E4M3) format (maximum finite value 448).
- BLAKE3 cryptographic hash function
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference.
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization.
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models.
- Pearl PoUW protocol whitepaper.
Copyright
Copyright and related rights waived via CC0.