Towards Generalizable 3D Anomaly Detection
via Relational Inconsistency Modeling

NeurIPS 2026

Kyung Hee University, Republic of Korea
*Equal Contribution    †Corresponding Author

GRIM learns an explicit, category-agnostic defect criterion
by modeling anomalies as relational inconsistencies among neighboring 3D structures.

Abstract

3D anomaly detection identifies defective regions in point clouds, but most existing methods are normality-centered: they learn normal data distributions and treat deviations as anomalies. This creates ambiguous decision boundaries, especially in unified and cross-domain settings where diverse normal distributions blur the boundary between normal and abnormal samples. We propose GRIM, a relational inconsistency modeling framework that instead characterizes defects as violations of geometric consistency among neighboring structures. GRIM learns category-agnostic defect cues from controlled pseudo-anomalies and contains two main components: Edge-aware Graph Refinement (EGR) for relation-aware local feature learning and Cluster-Deviation Modeling (CDM) for identifying regions that are incompatible with their structural peer groups.

Motivation: From Normality-Centered Detection
to Relational Defect Learning

Prior 3D anomaly detection methods mainly learn the distribution of normal samples and classify deviations as anomalies. However, this strategy does not explicitly define what makes a region defective. As a result, rare-but-valid normal regions can be over-penalized, while real defects that still resemble normal geometry can be missed.

Figure 1. Comparison of previous normality-centered methods and the GRIM framework.

Figure 1. Comparison between previous normality-centered 3D anomaly detection methods and our relation-based defect learning framework. GRIM uses pseudo-anomalies as controlled violations of geometric relations and learns a more discriminative defect criterion, which reduces both false positives and false negatives.

The left side illustrates the typical failure mode of previous approaches: the decision boundary is defined only with respect to the normal data manifold, so the model can confuse unusual normal structures with anomalies and fail to isolate real anomalous regions. The right side shows the key shift in GRIM: instead of only learning "what normal looks like," we explicitly train the model to recognize relation-violating structures through pseudo-anomaly generation.

Key idea. GRIM does not try to enumerate every possible real defect appearance. Instead, it teaches the model a transferable rule: defects break geometric consistency among neighboring or structurally similar regions.

Why Previous Unified 3D-AD Methods Fail

To analyze the weakness of normality-centered unified 3D anomaly detection, we compare GRIM with MC3D-AD using two complementary perspectives: false positives at fixed true positive rates and false negatives at fixed true negative rates. This analysis highlights how ambiguous decision boundaries lead to both false alarms and missed detections.

Figure 2. False positive and false negative comparison between MC3D-AD and GRIM.

Figure 2. Comparison of False Positive Rate (FPR) and False Negative Rate (FNR) between the unified baseline MC3D-AD and GRIM. The blue curves correspond to GRIM and the gray curves to MC3D-AD.

GRIM consistently yields lower FPR and FNR in both in-domain and cross-domain settings. This means that GRIM not only reduces false alarms on normal regions, but also improves sensitivity to real defects. The gain is especially clear in the cross-domain setting, showing that relation-based defect learning generalizes better than simply modeling category-dependent normality.

Interpretation. A better 3D anomaly detector should suppress false responses on valid structures while still detecting subtle anomalies. Figure 2 shows that GRIM achieves both simultaneously.

Proposed Method: GRIM

GRIM is designed around a simple principle: a model should learn why a region is defective, not only whether it deviates from normality. The overall framework is illustrated below.

Figure 3. Overall architecture of GRIM.

Figure 3. Overview of GRIM. Starting from a normal point cloud, GRIM generates pseudo-anomalies, extracts local features, refines them with Edge-aware Graph Refinement (EGR), and measures cluster-relative inconsistency with Cluster-Deviation Modeling (CDM) to produce anomaly scores.

1. Pseudo-Anomaly Generation

GRIM synthesizes localized bulge, sink, and hole perturbations on normal point clouds. These perturbations are controlled relation-violating samples rather than attempts to exhaustively simulate all real defects. Their role is to provide explicit supervision for defect learning.

2. Edge-aware Graph Refinement (EGR)

Each sample is partitioned into local groups and represented as a graph. Node features encode local regions, while edges describe relative geometric relationships such as position, normal, spread, and curvature. EGR refines the node features through relation-aware message passing.

3. Cluster-Deviation Modeling (CDM)

EGR-refined node features are clustered into structurally similar groups. GRIM computes how much each region deviates from its cluster center and injects this cluster-relative deviation into the anomaly classifier. This allows the model to detect regions that are incompatible with their structural peer group.

4. Test-Time Inference

At inference time, a test point cloud passes through the same feature extractor, EGR, and CDM pipeline. The final classifier predicts anomaly scores that can be used for both object-level detection and point-level localization.

Quantitative Results

97.4
Anomaly-ShapeNet
O-AUROC
94.5
Anomaly-ShapeNet
P-AUROC
87.2
Real3D-AD
O-AUROC
91.3
Real3D-AD
P-AUROC

In-domain Comparison

Mean AUROC (%) on Anomaly-ShapeNet and Real3D-AD. O-AUROC and P-AUROC denote object-level detection and point-level localization performance, respectively.

Method Anomaly-ShapeNet Real3D-AD
O-AUROC ↑ P-AUROC ↑ O-AUROC ↑ P-AUROC ↑
Category-Specific
BTF (Raw) (CVPR'23)49.355.063.557.1
BTF (FPFH)52.862.860.373.3
M3DM (CVPR'23)55.261.659.462.0
PatchCore (FPFH) (CVPR'22)56.858.059.368.2
PatchCore (PointMAE)56.257.759.462.0
CPMF (PR'24)55.957.358.675.8
IMRNet (CVPR'24)66.165.072.5-
Reg3D-AD (NeurIPS'23)57.266.870.470.5
Group3AD (ACM MM'24)--75.173.5
R3D-AD (ECCV'24)74.9-73.4-
ISMP (AAAI'25)-69.176.783.6
PO3AD (CVPR'25)83.989.876.565.0
PASDF (ICCV'25)90.089.780.274.5
Reg2Inv (NeurIPS'25)86.188.278.087.8
Unified
MC3D-AD (IJCAI'25)84.275.978.276.8
GRIM (Ours) 97.4 94.5 87.2 91.3

GRIM is trained as one unified model across categories, without category labels or category-specific parameters.


Cross-domain Generalization

Cross-domain evaluation directly applies the source-trained model to the target dataset without fine-tuning or domain-specific adaptation. Parenthesized values are the corresponding in-domain performance and Δ denotes the in-domain-to-cross-domain performance gap.

Method Anomaly-ShapeNet → Real3D-AD Real3D-AD → Anomaly-ShapeNet
O-AUROC ↑ Δ ↓ P-AUROC ↑ Δ ↓ O-AUROC ↑ Δ ↓ P-AUROC ↑ Δ ↓
MC3D-AD 55.4 (78.2)22.838.3 (76.8)38.5 78.3 (84.2)5.948.8 (75.9)27.1
GRIM (Ours) 83.6 (87.2)3.689.6 (91.3)1.7 91.7 (97.4)5.778.9 (94.5)15.6

Qualitative Results

Figure 4 compares GRIM and MC3D-AD under both in-domain and cross-domain settings. The results show that GRIM localizes defects more precisely and avoids the excessive false responses that often appear in normality-centered reconstruction-based baselines.

Figure 4. Qualitative comparison between MC3D-AD and GRIM.

Figure 4. Qualitative comparison between MC3D-AD and GRIM. The left half shows in-domain results (trained and tested on Anomaly-ShapeNet), and the right half shows cross-domain results (trained on Anomaly-ShapeNet and tested on Real3D-AD). Red regions indicate anomaly responses.

MC3D-AD often produces overly broad or misplaced anomaly activation maps, especially in cross-domain cases. In contrast, GRIM concentrates its responses on the actual defective regions and produces cleaner, more spatially precise localization. This qualitative behavior supports the main claim of the paper: learning a relation-based defect criterion improves both robustness and generalization.

Citation

@article{heo2026towards,
  title={Towards Generalizable 3D Anomaly Detection via Relational Inconsistency Modeling},
  author={Heo, KunHo and Kim, SuYeon and Lee, Hayoung and Oh, Chanse and Cho, MyeongAh},
  journal={arXiv preprint arXiv:2609.35059},
  year={2026}
}