When a sensor only sees a tracked object across a handful of points per pass, the cheapest way to recover most recognisable structure is to stack those passes and ignore the order they arrived in. Add a sequence model on top, and nothing changes.
That is the working heuristic in a practitioner ablation on the public RadarScenes dataset, where a single scan returns about 2.9 points on a tracked object. The brunopinto900 report puts a single-scan DeepReflecs baseline at 0.7370 macro F1, jumping to 0.8613 after pooling 20 scans into a bag, a +0.1243 lift that ignores scan order. A causal GRU on frozen per-scan embeddings adds +0.0282 to 0.8895; a fused GRU plus pooled-embedding variant registers +0.0002, which the report's author calls noise. Larger GRUs, a Transformer, a state-space model, and point-level self-attention all land in a 0.86 to 0.89 band, and the spread between the cheapest baseline and the fanciest stack is dwarfed by the lift simple aggregation already paid for.
The reusable rule: at this sparsity, point count on the target is the binding constraint, not modelling of scan order. Sequence architectures are icing on a pooling layer doing the heavy lifting. The counterweight is that pooling and the GRU differ in training and representation, so this is a ratio across one dataset and the brunopinto900 project's retraining runs; the project README warns the +0.0002 fusion increment sits inside a roughly ±0.002 macro-F1 retrain band. For sensor and ML practitioners: budget for temporal aggregation first, treat order-modelling as a small extra, and resist a complexity premium for a delta the cheap baseline has already captured.
Reported by Sky for Type0, from Multi scan radar object classification on RadarScenes [P]. Read the original: reddit.com