Batched two-body propagation
Propagate independent two-body trajectories with RK4.
24.1× on batch.64k · NVIDIA A100-SXM4-40GB · end to end · validation matched
Inspect public receipt ↗Raddle Forge coordinates and records validated optimization work. A coding agent or engineer supplies candidates; Forge gates them on correctness before benchmarking.
uvx raddle agent initUse the Raddle acceleration skill in this repository. Profile the existing project and identify one measured repetitive numerical bottleneck. Preserve the reference implementation and existing correctness tests. If a compatible Raddle accelerator exists, integrate it using the smallest safe change. Validate against the reference and benchmark the same representative workload before and after. Keep the integration only if validation passes and the measured result is worthwhile. If no accelerator matches, return a Raddle Forge candidate report instead of inventing one. Return the exact workload, baseline timing, accelerated timing, speedup, validation result, provenance, changed files, and reproduction commands.
from project.orbits import propagate_rk4 results = []for study in studies: result = propagate_rk4( study.initial_states, study.mu, study.dt, study.steps ) results.append(result)from raddle.orbit import PropagationInputfrom raddle.cupy_backend import accelerated, load backend = load()results = []for study in studies: inputs = PropagationInput( study.initial_states, study.mu, study.dt, study.steps ) results.append(accelerated(inputs, backend))65,536 trajectories · NVIDIA A100 · end-to-end · ✓ validation matched
Representative project integration. The measurement uses the public batch.64k case, comparing its numpy.vectorized and cupy.vectorized paths.
NVIDIA A100-SXM4-40GB · Raddle 0.3.0 · Python 3.12.12 · NumPy 2.5.3 · CuPy 14.2.0 · transfers included · setup excluded
Start with your agent. Inspect the public contracts when a candidate matches.
Propagate independent two-body trajectories with RK4.
24.1× on batch.64k · NVIDIA A100-SXM4-40GB · end to end · validation matched
Inspect public receipt ↗Compute deterministic bootstrap means and confidence intervals.
168.9× on bootstrap.64k · NVIDIA A100-SXM4-40GB · end to end · validation matched
Inspect public receipt ↗Forge coordinates and records candidate validation and benchmark evidence. External coding agents or engineers supply the candidate implementations; each speedup below includes its stated execution boundary.
NumPy reference → fused CuPy · host input, transfers, compute, and host output included · 94.75 ms → 4.10 ms. Transfer contribution: 25.4% of candidate time.
NPBench case npbench.jacobi_2d.m · git:f2d7f27d87ba4881991e62c18a607aa6ba52b260Inspect Jacobi campaign ↗Pinned PyTorch reference → fused GPU candidate · 8.98 ms → 6.57 ms · upstream correctness contract passed · A100.
kernelbench.l2.1_conv2d_relu_biasadd · ref git:423217d9fda91e0c2d67e4a43bf62f96f6d104f1external.kernelbench.l2.1_conv2d_relu_biasadd · artifact readback verifiedThese campaign histories are distinct from checked benchmark receipts.
Bring a measured workload ↗Future capability · managed execution is not available yet.
Future capability · customer-controlled cloud, HPC, and on-prem execution are not available yet.