WebGPU Shader & Inference Latency Profiler
Benchmark local GPU compute capabilities using the native WebGPU API. Measure WGSL compute shader compilation, parallel GEMM matrix throughput (GFLOPS), and VRAM transfer bandwidth offline.
[ 1. HARDWARE ADAPTER & GPU LIMITS ]
GPU Vendor
Detecting...
Architecture / Device
Detecting...
Max Storage Buffer
Detecting...
Max Workgroup Size (X,Y,Z)
Detecting...
ADVANCED WEBGPU EXTENSIONS:
shader-f16 (FP16 Tensors)
subgroups (SIMD Matrix)
timestamp-query
float32-filterable
[ 2. BENCHMARK CONTROLS ]
Parallel tiled matrix multiplication compute shader executed purely on GPU.
[ 3. MEASURED PERFORMANCE METRICS ]
Compute Throughput
--
GEMM Single-Precision GFLOPS
Shader Compile Latency
--
WGSL Pipeline Creation (ms)
VRAM Readback Speed
--
GPU-to-CPU Buffer Copy (GB/s)
Inference Class
--
Edge AI Suitability
[ PROJECTED LOCAL BROWSER LLM TOKENS/SEC ]
| Model Architecture | Est. KV Cache VRAM | Projected Speed |
|---|---|---|
| Qwen 2.5 0.5B (q4f16) | ~450 MB | -- tok/s |
| Llama 3.2 1B (q4f16) | ~880 MB | -- tok/s |
| Llama 3.2 3B (q4f16) | ~2.2 GB | -- tok/s |
| Mistral 7B (q4f16) | ~4.8 GB | -- tok/s |
[ 4. DIAGNOSTIC LOG & JSON REPORT ]
[ SYSTEM ] WebGPU Profiler initialized. Ready to execute WGSL matrix kernel.