GPU-Accelerated Adaptive Chirplet Transform (ACT)
Residual-based signal decomposition with a hybrid CPU–GPU implementation, evaluated across desktop, laptop, and embedded hardware.
This recorded example decomposes two seconds of EEG (512 samples at 256 Hz) from channel HB_1 in the Bitbrain sleep dataset. Each step adds a chirplet to the approximation. After 12 chirplets, the residual norm is 4.28% of the original signal norm, corresponding to about 0.18% residual energy. Playback illustrates precomputed results; it does not measure processing latency.
How It Works
The Problem
The Adaptive Chirplet Transform represents a non-stationary signal using a small set of time–frequency atoms. This work addresses divergent behavior in a prior implementation and the cost of searching a high-dimensional chirplet parameter space.
- Algorithmic instability. The per-step search for the best-fit chirplet happens over a high-dimensional, non-convex parameter space. Without careful numerical treatment, the decomposition can diverge, and later steps make the residual worse, not better.
- Computational cost. Dictionary search and generation can limit throughput, particularly as signal length, grid resolution, or channel count increases.
The Approach
Unit normalization and residual projection. Candidate atoms are normalized to unit L2 norm and projected onto the current residual. Subtracting the corresponding projection gives a non-increasing residual norm in exact arithmetic. The paper compares this implementation with the earlier method on EEG, EMG, and radar signals.
Hybrid CPU–GPU implementation. The CPU constructs the chirplet dictionary; the GPU performs batched projection and residual updates. Continuous parameter refinement uses BFGS. In the tested implementations, GPU-only dictionary generation added substantial overhead, so the hybrid design retains CPU generation and reuses the dictionary across epochs.
The dictionary is generated on the CPU and reused. Each updated residual feeds back into the GPU projection search; continuous parameter refinement is omitted from this simplified diagram.
Measured Runtime
Table IV of the paper reports three-second EEG epochs, decomposition order 10, on an RTX 5070 with an Intel i9-14900K and 64 GB RAM. Results are mean ± standard deviation across five runs, with 50 evaluated epochs per run and an initial warm-up epoch excluded.
| Implementation | Time per epoch | Overall run |
|---|---|---|
| CPU | 0.166 ± 0.023 s | 15.342 ± 0.238 s |
| CuPy hybrid | 0.025 ± 0.005 s | 8.372 ± 0.120 s |
| PyTorch hybrid | 0.022 ± 0.005 s | 7.952 ± 0.016 s |
Overall runtime includes dictionary generation and epoch processing. Per-epoch acceleration is approximately 6.6–7.5× using the rounded table values; the overall run is about 1.9× faster. Separate multi-channel experiments report a speedup rising from 3.94× at one channel to 8.22× at ten channels. These results describe the tested workloads, rather than a universal speedup.
Seeing It Work
For this example, the residual norm ratio decreases from 0.430 after the first chirplet to 0.0428 after the twelfth. Squaring these ratios gives residual energy fractions of 18.5% and 0.18%, respectively. This example illustrates reconstruction behavior; the benchmarks below quantify runtime separately.
Status & Citation
Under review at IEEE Open Journal of Signal Processing (submitted July 2026). Full derivation, benchmarks, and multi-channel scalability results are in the paper — N. Kumar, S. Mann, “Toward a Stable and Deployable Adaptive Chirplet Transform: Residual Projection, Hybrid GPU Acceleration, and Multi-Channel Scalability”, arXiv:2607.16629, 2026.
BibTeX Citation
@article{kumar2026chirplet,
title={Toward a Stable and Deployable Adaptive Chirplet Transform: Residual Projection, Hybrid GPU Acceleration, and Multi-Channel Scalability},
author={Kumar, Nishant and Mann, Steve},
journal={arXiv preprint arXiv:2607.16629},
year={2026}
}