Research notes

What Sora and Veo videos carry, and what survives a re-upload

Content Credentials, device tags and track names in real sample videos, and what a one-line FFmpeg copy kept and removed.

Updated 24 September 2026

We went looking for a Sora or Veo video with its original bytes intact, to show exactly what an AI video file carries. We didn’t find one. The Sora and Veo pages we checked don’t offer their clips as files, and a copy from a stranger’s re-upload would tell us about their software, not OpenAI’s or Google’s. So this page works from the five public test videos that do carry Content Credentials, plus two that don’t, and from what OpenAI, Google, TikTok and YouTube say in their own documents.

What we can show is how little it takes to lose the metadata. A plain copy of a video through FFmpeg, one command and no re-encoding, removed the Content Credentials from all four signed files we tried.

What the makers say they put in

OpenAI’s launch post for Sora says every output carries a visible watermark and that “all Sora videos also embed C2PA metadata”, the signed record better known as Content Credentials. Google’s Veo documentation for developers says videos made by Veo “are watermarked using SynthID”; that page doesn’t mention C2PA at all. Those are two different things, and they sit in different places:

  • Content Credentials are a block of data inside the file. In an MP4 or MOV they live in their own box near the start, which our checker reads and lists.
  • SynthID is an invisible watermark in the picture and sound. No file reader sees it, ours included.
  • The Sora logo is part of the picture: you see it when the video plays, and it has nothing to do with the file’s metadata.

Whether a site then shows an “AI” label is a third thing again, decided by that site. The three layers fail in different ways, so they need checking separately.

Seven public sample videos, opened

These come from the test files of the C2PA’s own software, the C2PA public test set and the IPTC’s video samples. We opened each with our AI video metadata checker and cross-checked every one with ExifTool 13.55.

Public sample videos read by the WipeTheAI video metadata checker in Chromium 153 and cross-checked with ExifTool 13.55, 24 September 2026 (n=7 files). Box size is the Content Credentials box alone
FileSizeCredentials boxSigned by (certificate name)What else the file names
video1.mp4 (C2PA SDK tests)828,571 bytes30,662 bytes, 2 recordsC2PA Test Signing CertXMP edit history naming Adobe Premiere Rush; track names from Mainconcept
dashinit.mp4 (C2PA SDK tests, streaming fragment)4,765 bytes3,591 bytes, 1 recordA test publisher certificateFFmpeg’s encoder tag, Lavf58.31.104
Truepic zoetrope clip (C2PA public test set)15,456,823 bytes30,192 bytes, 1 recordTruepic Lens SDK in Vision CameraAndroid version 13; the record holds camera EXIF and motion data
IPTC park bench, generator sample4,841,832 bytes16,151 bytes, 1 recordA test publisher certificateFFmpeg’s encoder tag; an XMP block
IPTC news agency sample20,114,238 bytes29,974 bytes, 1 recordIPTC C2PA Signer ToolSays “scanned from negative film”; author and title tags; a 172,790-byte thumbnail
IPTC park bench, original4,822,934 bytesNoneNot signedFFmpeg’s encoder tag only
c.mov, a Mac screen recording (C2PA SDK tests)973,831 bytesNoneNot signedMake Apple, model Mac15,9, software macOS 14.6.1, with the recording’s local time

Two things stood out. None of the five signed files says generative AI made it: only one carries a digital source type at all, and it says film scan. So these samples show the structure, not what Sora writes into it. And the credentials are small. In the Truepic clip the box is 30 KB of a 15 MB file, 0.2%, which is one reason an app can drop it without anyone noticing a change in size.

Every signed file also records a fingerprint of the picture and sound data (the c2pa.hash.bmff record, in versions 1 to 3). That is what makes the credentials fragile: trim or re-encode the video and the fingerprint no longer matches, even where the box survives.

What survived a copy through FFmpeg

We ran five of the files through FFmpeg 8.1 three ways: a straight copy with -c copy, which rewrites the container but keeps the picture and sound as they are; the same with -map_metadata -1, the flag most answers suggest for removing metadata from a video; and a full re-encode to H.264. Then we opened all fifteen results in the checker and in ExifTool.

video1.mp4, the IPTC park bench generator and news agency samples, the Truepic clip and c.mov through FFmpeg 8.1 on macOS, 24 September 2026; results read by the WipeTheAI checker and ExifTool 13.55 (n=5 files, 15 outputs)
What was in the fileStraight copy (-c copy)Copy with -map_metadata -1Re-encoded to H.264
Content Credentials box (4 files)Gone in 4 of 4Gone in 4 of 4Gone in 4 of 4
XMP block (3 files)Gone in 3 of 3Gone in 3 of 3Gone in 3 of 3
Apple make, model and software (c.mov)GoneGoneGone
Android version tag (Truepic clip)GoneGoneGone
Author, title and thumbnail tags (IPTC news sample)GoneGoneGone
Creation time in the file header (4 files)Reset to zero in 4 of 4Reset to zero in 4 of 4Reset to zero in 4 of 4
Track names such as “Core Media Video” (5 files)Kept in 5 of 5Replaced with “VideoHandler” in 5 of 5Kept in 5 of 5
New tag addedFFmpeg’s encoder, Lavf62.12.100FFmpeg’s encoder, Lavf62.12.100FFmpeg’s encoder, Lavf62.12.100

The straight copy removed more than we expected, and it isn’t a clean result either. The Mac recording lost its make, model and macOS version but kept the track name “Core Media Video”, the name Apple’s Core Media framework gives the tracks it writes. And every output gained a line saying FFmpeg made it. In all fifteen of our outputs, something still named a program that had saved the file.

One flag brings a lot back. Adding -movflags use_metadata_tags to the copy kept the Mac recording’s make, model, macOS version and local recording time, and the Truepic clip’s Android version. The Content Credentials still didn’t come through, because FFmpeg doesn’t carry that box over in any mode we tried.

If your aim is to remove metadata from your own video before sharing it, for the location or the phone model, the straight copy with -map_metadata -1 is the most thorough of the three FFmpeg commands we tested. Check the result rather than trusting the flag: drop it into the checker and see what is left.

Removing it without re-encoding, and proving nothing else changed

FFmpeg rebuilds the whole container, which is why it adds its own tag. Our video metadata remover does less: it cuts out the Content Credentials box, the XMP block and the tag boxes, blanks the created and modified times in the header, and copies every other byte as it was. Cutting boxes that sit before the picture and sound moves that data, so every pointer to it (the chunk offset tables in each track, and in streaming files the fragment and index offsets) has to be recalculated. Get one wrong and the video breaks, so we checked each output three ways: FFmpeg’s framemd5 hash of every packet as stored, the same hash of every decoded frame, and whether Chromium would open the file.

Five public sample videos cleaned by the WipeTheAI video metadata remover in Chromium 153, 24 September 2026. Packets and frames compared with FFmpeg 8.1 framemd5 before and after; leftovers read with ExifTool 13.55 (n=5 files)
FileTaken outBytes savedPackets and decoded frames identicalOpens in Chromium
video1.mp4Credentials box (30,662 bytes), XMP (4,562), timecode tags (58)35,282154 of 154 packets, 154 of 154 framesYes, 720 × 1280, 2.0 s
Truepic zoetrope clipCredentials box (30,192 bytes), Android version tag (118)30,310549 of 549, 549 of 549Yes, 1920 × 1080, 7.4 s
IPTC news agency sampleCredentials box (29,974 bytes), tags with the thumbnail (172,934), XMP (9,497)212,405797 of 797, 795 of 795Yes, 1920 × 1080, 11.2 s
IPTC park bench, generator sampleCredentials box (16,151 bytes), XMP (2,747), encoder tag (98)18,996523 of 523, 522 of 522Yes, 1920 × 1080, 7.2 s
c.mov, Mac screen recordingApple make, model, macOS version, pixel density and time tags (4,609 bytes)4,609374 of 374, 310 of 310Yes, 2750 × 1834, 5.2 s

One result says more than the table. The signed park bench sample and its unsigned original, cleaned separately, came out as the same file, byte for byte (same SHA-256). Everything the signing step had added was gone, and nothing else about the file had moved.

ExifTool found two things left in every output: the blanked dates, and the track names such as “Core Media Video”, which sit inside boxes a player needs, so we keep them. Unlike the FFmpeg copies, no output gained a tag naming the program that saved it. The remover writes nothing of its own into the file.

We also cleaned files laid out the ways streaming services write them. FFmpeg made three fragmented copies of the park bench original (with absolute fragment offsets, with offsets counted from each fragment, and with a segment index), plus a copy with the header at the front, and a 603 MB file made by looping the Truepic clip 41 times. Every packet matched in all five. On the 603 MB file the report was on screen 28 milliseconds after the file was chosen, and the page’s JavaScript memory stayed at 16 MB, because only the header is read; the picture and sound are copied straight from the original file into the download.

To run the same check on your own file, hash every packet of both versions and compare the two lists (the first columns hold timings, the last one the hash):

ffmpeg -i original.mp4 -map 0 -c copy -f framemd5 before.txt

ffmpeg -i original-clean.mp4 -map 0 -c copy -f framemd5 after.txt

Some layouts the remover refuses rather than guessing: a movie whose media sits in another file, an old compressed QuickTime header, and streaming files where a fragment’s data position depends on the one before it. It leaves the file untouched and says why. And it only reaches the metadata layer. A SynthID watermark in the picture, the visible Sora logo and any label a site adds are all out of its reach.

What happens when you upload

We didn’t upload these files anywhere, so this part is only what the platforms say. TikTok wrote in May 2024 that it reads Content Credentials on upload to label AI content made elsewhere, and that it attaches its own credentials to TikTok videos, which stay on them when downloaded. YouTube’s “Captured with a camera” note appears only for videos recorded with tools supporting C2PA 2.1 or later and not edited afterwards; YouTube lists saving the file somewhere that doesn’t support C2PA as one way to lose it.

Put next to our FFmpeg result, the pattern is plain. The credentials survive only along a chain of software that chooses to keep them, and a site that labels from credentials can only see them if nothing in between re-saved the file. What a label is based on otherwise, whether a detector for invisible watermarks, what the uploader ticked or something else, is the site’s decision and isn’t in the file.

For your own videos, the video metadata remover does the job described above. For images, the Wipe tool removes Content Credentials and AI metadata without changing a pixel, and this guide covers how that works for pictures.

Sources

  1. OpenAI: Launching Sora responsibly
  2. Google AI for Developers: Generate videos with Veo (SynthID note)
  3. TikTok: Partnering with our industry to advance AI transparency and literacy (9 May 2024)
  4. YouTube Help: “Captured with a camera” disclosure
  5. C2PA SDK test videos (video1.mp4, dashinit.mp4, c.mov)
  6. C2PA public test files (Truepic zoetrope clip)
  7. IPTC Video Metadata Hub: C2PA sample files
  8. FFmpeg documentation: the mov/mp4 muxer and its movflags
  9. FFmpeg documentation: the framemd5 muxer (per-packet hashes)