跳转到内容

TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers

TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers

Section titled “TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers”

⚠️ AI 生成 · 建议对照原文 本页为自动整理的学习笔记;关键数据与引用如需引用,请回查 PDF / 官方版本。

学习档位 自动卡

类型 文献 · 更新 2026-07-19

所属 Query-based 检测与集合预测

  • topic: foundations
  • sources: arxiv, openalex
  • retrieved_at: 2026-07-20
  • query: End-to-End Object Detection with Transformers DETR
  • arxiv: 2201.05047
  • doi: 10.1109/TPAMI.2022.3223955
  • score_total: 64
  • suggested_tier: recent

(no prose relevance explanation — numeric score only or HTTP source)

(no snippet evidence in candidate pool)

AI analysis (heuristic fallback, needs-source-verification)

Section titled “AI analysis (heuristic fallback, needs-source-verification)”

Status: metadata_only · Source: metadata

Metadata-only card for TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers.

生成:2026-07-21 · 来源条数 0 · 模型 heuristic · 需人工核验数字

围绕「TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers」的核心问题与动机(待结合全文核验)。

  • 见原文方法章节;以下为基于摘要/摘录的要点提示。
  • 本地摘录暂缺。

与相近工作的关系待核验;请对照 related work。

  • 勿仅凭摘要推断未给出的数值指标。
  1. 这篇工作的输入/输出表示是什么?(TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers)
  2. 训练目标与评测协议各是什么?
  3. 主要失败模式或局限是什么?
  • (本次未抓取到白名单二次解读页)
flowchart LR
A["输入 / 观测"] --> B["表示 / 编码"]
B --> C["推理 / 解码"]
C --> D["输出 / 动作或检测"]
%% method sketch for: TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers

方法结构示意(重绘;细节以原论文为准,待 PDF 核验)。

展开英文 Paper Card / AI deep analysis
Field Content
Year 2022
Authors Qianyu Zhou, Xiangtai Li, Lu H, Yibo Yang, Guangliang Cheng, Yunhai Tong, Lizhuang Ma, Dacheng Tao
arXiv
DOI 10.1109/tpami.2022.3223955
Topics foundations
展开 Extract / Selections / Local assets
  • foundations: tier=watch rank=2 score=58 — auto refresh 2026-07-19 sources=openalex
(no PDF text available; metadata-only card)