Gallery inside!
Research

DEST for 3D Object Detection: Better Accuracy, Added Latency

DEST improves indoor 3D object detection by updating scene and object features together. Explore original diagrams, benchmark gains and the latency tradeoff.

6

Figures open at full size. Wide tables scroll sideways.

A robot or an augmented-reality system needs more than a cloud of depth measurements. It needs to identify objects and locate their boundaries: where the table ends, which points belong to a chair, and what occupies the surrounding space.

DEST studies a specific bottleneck in that process. Some 3D detectors repeatedly refine their object predictions while reusing the same underlying scene features. The researchers ask whether updating both representations together would let later processing layers do more useful work.

The answer is promising for indoor perception: detection accuracy improves on two benchmarks, but inference takes longer. For a team considering the architecture, the decision is whether those gains survive its own scenes, hardware and response-time requirements.

The Core Insight: Update the Scene as Well as the Objects

A point cloud represents a scene as many spatial samples. A detector first turns those points into features. It then maintains a smaller set of object candidates—often called queries—that gather information from the scene and become predicted boxes with category labels.

In the transformer-based baselines examined by the paper, the queries change from one decoder layer to the next, but the scene features supplied to those layers remain fixed. Think of repeatedly revising an interpretation while consulting an unchanged set of notes. Later revisions may add little if the notes never incorporate what earlier reasoning discovered.

Chuxin Wang and colleagues propose an interactive state space model, the central component of DEST. It passes updated scene features and updated object states into subsequent layers. The original comparison makes that change visible:

Original DEST Figure 1 comparing a transformer decoder with fixed scene features against a decoder that updates scene and object features together; adjacent charts show larger gains from later layers in the tested DEST variants.
Wang and colleagues, original Figure 1. The lower architecture passes updated scene information forward alongside the object candidates. The charts concern gains within stages of these models, not a multiplication of overall detection accuracy. Original figure and argument.

The important change is the information flow. DEST does not simply remove attention, and “state” does not mean that it learns continuously from a live video feed. In this paper, the sequence being processed consists of spatial points within a scene. Object tracking across time is a separate problem.

How the Detector Works

The method begins with an encoder that extracts scene features and a sampling step that selects initial object candidates. DEST focuses on the decoder that refines those candidates.

Each candidate acts as a state that accumulates relevant evidence from scene points. Crucially, different candidates must attend to different regions. A candidate near a chair should not update itself in exactly the same way as one near a table.

To make that possible, DEST predicts a box from each initial state and uses the spatial relationships between scene points and the box’s corners to help determine how information is incorporated. Scene features also influence that decision, helping distinguish useful object information from background. The model produces updated object states and updated scene features, so the next layer receives both.

Four supporting choices make this work with point clouds:

  • Give the points an order. A state space model processes a sequence, while a point cloud has no natural reading order. DEST sorts points using a space-filling curve that tends to keep nearby locations close in the sequence.
  • Read in both directions. A single forward pass would give later points access to earlier information without the reverse. Forward and backward scans let information travel both ways.
  • Let object candidates interact. DEST retains attention between object states to model relationships among objects. It is a hybrid design.
  • Refine the resulting features. A gated feed-forward network controls which feature information passes through; local convolutions also help capture nearby structure.

The ordering step is easier to understand in the paper’s appendix illustration:

Original appendix diagram showing two Hilbert curve axis orders and the different point sequences they produce from the same two-dimensional point set.
Original Figure 5, Wang and colleagues. The same points can produce different sequences when coordinate priority changes. The paper uses this two-dimensional illustration to explain its three-dimensional ordering strategy. Original appendix figure.

A single ordering cannot preserve every neighborhood relationship. DEST therefore uses different coordinate orders across decoder layers, giving the same scene several traversal patterns. In the appendix experiments, using six orders produces a higher detection score than repeating one order. Reading in both directions also beats a one-direction scan. The gain depends on adapting the sequence model to spatial data, rather than inserting a generic sequence block and hoping it works.

What the Research Actually Shows

The experiments use ScanNet V2 and SUN RGB-D, two indoor-scene benchmarks. ScanNet evaluation covers 18 object categories; the SUN RGB-D setup follows prior work in evaluating ten common categories. These are annotated scene datasets, not tests of warehouse throughput or robot safety.

The detector predicts a category and a 3D box for each object. The main comparison below uses AP50: average precision across categories, with a detection required to meet a 0.5 overlap threshold between its predicted box and the reference box. A higher score reflects a better detection result under that rule. It should not be read as the percentage of all objects a deployed system will identify correctly.

The researchers train models five times and evaluate each trained model five times. Their main results table reports both the best result and the average across those trials. Looking at the averages keeps the comparison from depending on a single favorable run:

  • On ScanNet V2, the smaller GroupFree baseline rises from 48.5 to 52.7 with DEST. The larger GroupFree variant rises from 51.8 to 56.8.
  • On SUN RGB-D, the smaller GroupFree model rises from 44.4 to 47.6.
  • With the stronger, color-using VDETR baseline, the averages rise from 64.5 to 66.2 on ScanNet and 49.7 to 50.9 on SUN RGB-D, without test-time augmentation.

These matched comparisons are more informative than treating every row as a level playing field. GroupFree uses point positions, while the VDETR configuration also uses color and a different encoder. The paper’s highest scores use test-time augmentation—additional transformed inputs at evaluation—which is another condition to preserve when comparing results.

A Generic State Space Block Was Not Enough

The component experiment is particularly revealing. Replacing the decoder with a more basic state space design initially performs worse than GroupFree. Ordering the points and adding bidirectional processing help, but the large recovery comes when the model uses DEST’s interactive, spatially informed state updates.

Original Table 2 showing GroupFree at AP50 48.5, a basic state space baseline at 41.6, and successive additions reaching 52.7 with the complete DEST design.
Original Table 2, Wang and colleagues. Rows add components cumulatively. The “Baseline” row is the basic state space replacement, not GroupFree. Results are averages from the ScanNet component experiments. Original component analysis.

There is also a direct check on the paper’s central idea. When DEST keeps the scene features fixed, its AP50 falls to 49.3; updating them produces 52.7. This supports the value of refreshing scene information. However, the complete detector also changes other components and adds training supervision for scene points, so the full improvement should not be attributed to one architectural choice alone.

The Accuracy Gain Has a Runtime Cost

The state space formulation is intended to make scene interaction efficient as the point sequence grows. That theoretical scaling property does not mean the finished detector is faster than the baseline.

On the paper’s Tesla V100 measurements, the smaller GroupFree model moves from 21 to 34 milliseconds, the larger version from 68 to 98 milliseconds, and the VDETR comparison from 238 to 263 milliseconds. Model parameter counts also increase.

Original Table 5 showing DEST increases latency and parameter counts relative to each GroupFree and VDETR baseline while improving the reported detection scores.
Original Table 5, Wang and colleagues; Tesla V100 inference measurements. This table pairs latency with the paper’s higher-score comparisons; the VDETR accuracy entries correspond to its test-time-augmentation results. Original speed comparison and conditions.

For a system with a tight response budget, 13 extra milliseconds in one stage can matter. For an offline scene-analysis workflow, the accuracy improvement may be worth considerably more. These measurements establish a tradeoff on the tested GPU, not performance on an edge device or an entire robotics pipeline.

Where It Could Help—and Where It Still Fails

Indoor mapping, scene understanding and AR object placement are plausible application areas when better boxes and category predictions justify extra computation. A team should first check whether its actual failure resembles the one DEST addresses: object candidates are present, but repeated refinement fails to use enough scene context.

The limitations appendix identifies a different failure that the decoder cannot reliably repair: some objects lack initial candidate points. If an object never gets a useful starting candidate, better refinement may still miss it. Candidate sampling is left for further work.

That is why a deployment evaluation should inspect missed objects by category and scene condition, rather than stopping at one aggregate score. It should also include the complete path from sensor input to the action or overlay that consumes the detections. The paper does not demonstrate industrial picking gains, continuously adapting perception, or transfer to object tracking.

Implementation Frameworks

Use the official GroupFree implementation to establish a reproducible baseline. Its repository provides model, data-preparation, training and evaluation code for ScanNet and SUN RGB-D. It is a historical research stack with older environment requirements, so begin in an isolated environment and verify that you can reproduce the baseline before modifying the decoder.

Use DEST’s method and appendix as the implementation specification. The paper describes the bidirectional scan in Algorithm 1, the ordering procedure and dataset-specific training settings. This is substantial model engineering, not a switch in a general-purpose agent framework. We have inspected the method; we have not reproduced its training results or established a ready-to-install DEST package here.

For a first comparison, keep the encoder, input sampling, candidate count, dataset split and evaluation procedure consistent. Add the interactive decoder and its specified training changes, then repeat evaluation across runs. Track detection quality, missed-object cases, memory and latency on the hardware you intend to use. If the result depends on color or test-time augmentation, carry those requirements into the product assessment.

For teams without capacity to reproduce that experiment, the immediate action is narrower: audit whether current misses come from candidate selection, scene representation or later refinement. DEST is a promising answer to one of those problems, and a poor reason to replace a functioning perception stack without that diagnosis.

TechClarity’s View

DEST’s contribution is more interesting than a claim that state space models are faster. It shows how updating scene and object representations together can improve a detector that has reached diminishing returns from repeated query refinement.

The component tests strengthen that argument, while the latency table defines its cost. We would consider it for accuracy-sensitive indoor perception with room in the compute budget. For strict real-time or unfamiliar sensor environments, adoption should depend on a matched local evaluation—not the architecture’s name or its historical benchmark position.

Original Research

State Space Model Meets Transformer: A New Paradigm for 3D Object Detection, by Chuxin Wang, Wenfei Yang, Xiang Liu and Tianzhu Zhang. Published as an ICLR 2025 conference paper; this article uses arXiv version 2, dated March 19, 2025. Original diagrams and tables are attributed to the authors and included beside the findings they explain.

Related research

Tags:
Author
TechClarity Analyst Team
September 27, 2026