RADDLE / VALIDATED ACCELERATION

Keep the calculation.
Lose the wait.

Raddle Forge coordinates and records validated optimization work. A coding agent or engineer supplies candidates; Forge gates them on correctness before benchmarking.

01 / INITIALIZE IN YOUR PROJECT
uvx raddle agent init
02 / GIVE YOUR CODING AGENT THE PROMPT
Paste the public kickoff prompt into your agent.
Read the prompt
Use the Raddle acceleration skill in this repository.

Profile the existing project and identify one measured repetitive numerical bottleneck. Preserve the reference implementation and existing correctness tests.

If a compatible Raddle accelerator exists, integrate it using the smallest safe change. Validate against the reference and benchmark the same representative workload before and after. Keep the integration only if validation passes and the measured result is worthwhile.

If no accelerator matches, return a Raddle Forge candidate report instead of inventing one.

Return the exact workload, baseline timing, accelerated timing, speedup, validation result, provenance, changed files, and reproduction commands.
BEFORE → RADDLE

Same calculation.
Measured difference.

orbit_studies.pyPROJECT / CODING AGENT OR ENGINEER
REPRESENTATIVE PROJECT CODE
from project.orbits import propagate_rk4 results = []for study in studies: result = propagate_rk4( study.initial_states, study.mu, study.dt, study.steps ) results.append(result)
from raddle.orbit import PropagationInputfrom raddle.cupy_backend import accelerated, load backend = load()results = []for study in studies: inputs = PropagationInput( study.initial_states, study.mu, study.dt, study.steps ) results.append(accelerated(inputs, backend))
BASELINE 2.94 sACCELERATED 121.9 ms
24.1×faster on this workload

65,536 trajectories · NVIDIA A100 · end-to-end · ✓ validation matched

Representative project integration. The measurement uses the public batch.64k case, comparing its numpy.vectorized and cupy.vectorized paths.

Inspect receipt

NVIDIA A100-SXM4-40GB · Raddle 0.3.0 · Python 3.12.12 · NumPy 2.5.3 · CuPy 14.2.0 · transfers included · setup excluded

Public receipt ↗
PUBLIC ACCELERATORS

Built for real workloads.

Start with your agent. Inspect the public contracts when a candidate matches.

orbit.two_body_rk4 / v0.2.0

Batched two-body propagation

Propagate independent two-body trajectories with RK4.

numpy.vectorized · cupy.vectorized

24.1× on batch.64k · NVIDIA A100-SXM4-40GB · end to end · validation matched

Inspect public receipt ↗
stats.bootstrap / v0.1.0

Bootstrap resampling

Compute deterministic bootstrap means and confidence intervals.

numpy.vectorized · cupy.vectorized

168.9× on bootstrap.64k · NVIDIA A100-SXM4-40GB · end to end · validation matched

Inspect public receipt ↗
RADDLE FORGE

Validated results,
on real workloads.

Forge coordinates and records candidate validation and benchmark evidence. External coding agents or engineers supply the candidate implementations; each speedup below includes its stated execution boundary.

SCIENTIFIC NUMPY → GPU · END-TO-ENDNPBench Jacobi 2D · preset M23.11×

NumPy reference → fused CuPy · host input, transfers, compute, and host output included · 94.75 ms → 4.10 ms. Transfer contribution: 25.4% of candidate time.

NPBench case npbench.jacobi_2d.m · git:f2d7f27d87ba4881991e62c18a607aa6ba52b260Inspect Jacobi campaign ↗
OPTIMIZED GPU → GPU · A100KernelBench Conv2D + ReLU + BiasAdd1.37×

Pinned PyTorch reference → fused GPU candidate · 8.98 ms → 6.57 ms · upstream correctness contract passed · A100.

kernelbench.l2.1_conv2d_relu_biasadd · ref git:423217d9fda91e0c2d67e4a43bf62f96f6d104f1external.kernelbench.l2.1_conv2d_relu_biasadd · artifact readback verified
Inspect the SSAPy campaign lifecycle ↗

These campaign histories are distinct from checked benchmark receipts.

Bring a measured workload ↗
RUN

Raddle Run

Future capability · managed execution is not available yet.

DEPLOY

Raddle Deploy

Future capability · customer-controlled cloud, HPC, and on-prem execution are not available yet.