[논문리뷰] DINOv3: Self-supervised Learning for Vision at Scale (Meta AI, 2025)
7B ViT와 Gram anchoring으로 dense feature 붕괴를 잡고 frozen backbone만으로 종합 SOTA를 달성한 SSL foundation model
7B ViT와 Gram anchoring으로 dense feature 붕괴를 잡고 frozen backbone만으로 종합 SOTA를 달성한 SSL foundation model
curated LVD-142M 데이터와 개선된 iBOT 기반 recipe로 ViT-g(1.1B)를 학습하고 distillation로 전파해 OpenCLIP를 넘어선 범용 visual feature
label 없는 self-distillation로 ViT를 학습해 attention map에 segmentation이 창발하고 k-NN 분류 78.3%에 도달하는 DINO
noun phrase와 exemplar prompt로 concept의 모든 instance를 검출·분할·추적하는 promptable concept segmentation 모델 SAM 3 논문 리뷰
1M 합성 데이터셋과 monocular prior 적응, disparity transformer로 zero-shot stereo matching을 실현한 논문 리뷰
3DGS novel view 렌더링과 EfficientSAM zero-shot 분할로 재배치된 물체의 3D mask와 6D pose 변화를 추정하는 change detection 논문 리뷰
SE(3)-equivariant 표현 하나로 시간차 scan의 instance matching·registration·reconstruction을 함께 푸는 MoRE² 논문 리뷰
CAD model 유무와 무관하게 novel object의 6D pose estimation과 tracking을 하나의 framework로 통합한 FoundationPose 논문 리뷰
기준 view를 없애고 permutation-equivariant 구조로 affine-invariant pose와 scale-invariant point map을 예측하는 π³ 논문 리뷰
이미지와 선택적 geometric 입력에서 factored representation으로 metric 3D를 직접 회귀하는 통합 feed-forward 모델 MapAnything 논문 리뷰
camera·depth map·point map·3D point track을 단일 feed-forward transformer의 forward 한 번으로 예측하는 3D vision 통합 모델 VGGT 논문 리뷰
optical flow로 model-to-frame과 frame-to-frame registration을 결합해 unseen object의 6DoF pose를 정제·추적하는 GoTrack 논문 리뷰
기존 tracker들을 teacher로 삼아 real video에 pseudo-label을 달고 1,000배 적은 데이터로 SOTA를 넘는 단순화된 point tracker CoTracker3 논문 리뷰
track 간 attention과 support point로 여러 점을 jointly 추적해 occlusion에 강한 transformer 기반 point tracker CoTracker 논문 리뷰
depth와 ray map을 최소 prediction target으로 삼아 plain transformer 하나로 any-view geometry를 통일한 Depth Anything 3 논문 리뷰
실측 label을 전부 synthetic 이미지로 바꾸고 62M pseudo-labeled 실사 이미지로 증류해 정밀도와 robustness를 함께 잡은 Depth Anything V2 논문 리뷰
62M unlabeled 이미지 self-training과 DINOv2 feature alignment로 만든 zero-shot MDE foundation model, Depth Anything 논문 리뷰
matching을 3D 문제로 정식화해 DUSt3R에 local feature head와 InfoNCE loss를 더하고 Map-free localization에서 30%p 개선을 이룬 MASt3R 논문 리뷰
카메라 calibration도 pose도 없이 이미지 쌍에서 pointmap을 직접 회귀해, depth·대응·pose·3D 복원까지 기하 3D 비전 과제 전체를 하나의 모델로 푸는 DUSt3R 논문 리뷰
SAM 기반 ISM과 background token point matching의 PEM으로 zero-shot 6D pose를 구하는 SAM-6D 논문 리뷰