Visual Claim Forensics: A Phrase-Level Framework for Reasoning about Image Provenance
Abstract
Image-forensics systems can detect manipulation and localize suspicious regions, but a final image generally does not identify the source image or the edits that produced it. We use the claim accompanying an image to specify which visual evidence should be audited and introduce \emph{visual claim forensics}: phrase-level, set-valued inverse inference from a final image and claim. The framework compiles the claim into visually decidable predicates, fixes their grounding in the final image, and enumerates source--history pairs under an explicit bounded edit model. Rather than selecting a single speculative reconstruction, we retain all compatible histories and examine how the visual support for each predicate changes across them. Agreement between histories reveals source-invariant information, whereas disagreement makes the predicate’s provenance explicitly non-identifiable. The compatible histories can also indicate where a relevant edit may have occurred and provide concrete candidate-source witnesses. In a controlled CLEVR-based pilot, exact enumeration identifies 1,440 compatible histories, grouped into 856 phrase-conditioned classes. Although source ambiguity persists in every collision group, relational predicates that remain invariant across histories are still identifiable. The pilot is intentionally small and synthetic: it validates the representation and exact reasoning layer as a first step toward natural-image systems that combine forensic traces, provenance metadata, watermarks, or learned priors to rank constrained source hypotheses without hiding uncertainty.