Search Video By Content, Not Filename

Search Video By Content, Not Filename

By Connor Bower

You know the clip exists. You can picture it – the wide shot of the harbor, late light, the boat crossing frame left. What you cannot do is name it, because your camera named it, and your camera called it C0032.MP4.

This page covers what it actually takes to search a video library by content instead of by filename: why the filename is a dead end, what the three kinds of "content" in a video are, what a content index physically stores on your disk, and what that kind of search can and cannot retrieve.

The Short Answer

Searching video by content means searching an index built from the video itself rather than from its name. A visual index samples frames from every clip, converts each frame into a numeric vector, and compares your typed description against those vectors. You type "harbor at sunset" and get clips that look like that, whatever they are called.

A search results grid in VidFinder for the plain-language query "golden hour", showing clips whose filenames are camera codes

Why Filename Search Fails On Video

Filename search fails on video because the name was written by a camera, not by you. A Sony body names clips C0032.MP4. An iPhone writes IMG_4021.MOV. Spotlight and Finder index that name plus the file's metadata – duration, codec, resolution, dates – but never the picture inside it, so a search for "sunset" returns nothing.

The problem has three parts.

  • Camera-generated names carry no meaning. They are sequence numbers. Two clips shot four months apart on different continents can sit one digit away from each other.
  • Renaming does not scale. Naming 40 clips after a shoot is a chore you will do once. Naming 4,000 across three drives is a project nobody finishes.
  • Folders encode one dimension only. A clip filed under 2026-04 Lisbon Trip can be found by trip. At various points you will also want it as a drone shot, as golden-hour footage, and as a clip with nobody in it, and the folder has no way to say any of that.

What "Content" Actually Means Inside A Video

Video content comes in three separate layers: what the frames look like, what people say, and text printed on screen. Each layer needs its own index, built by different technology. A tool that reads visuals will not find a spoken sentence, and a transcript search will not find a wide shot of a harbor with nobody talking.

LayerWhat it indexesWhat it findsWhat it misses
VisualSampled frames, as image data"beach", "two people at a table", "close-up of hands", "empty street"Anything not visible in frame
SpeechAn audio transcriptA quoted line, a name someone said, a topic discussedSilent footage, B-roll, music-only clips
On-screen textCharacters read out of the pixelsA slide title, a lower third, a URL burned into the frameAnything not written on screen

Most footage libraries are dominated by the first layer. B-roll, drone shots, cutaways, and establishing shots have no dialogue and no on-screen text, so speech and text indexes have nothing to work with. For those libraries, only the visual index has anything to read.

VidFinder indexes the visual layer only. It has no speech transcription and no on-screen text recognition. If your footage is talking-head interviews and you need to find a quoted line, a transcript tool is the right answer, and we would rather tell you that than sell you the wrong index.

What A Content Index Actually Stores

A visual content index holds a short numeric summary of a handful of frames from each video, and no copy of the video itself. VidFinder samples eight frames per clip, converts each frame into a 512-number vector, and keeps those vectors in a local database – roughly 16 KB per video, or about 16 MB per 1,000 videos.

The numbers for VidFinder, rarely published elsewhere in this category:

  • Frames sampled per video: 8, taken at 5%, 16%, 27%, 39%, 50%, 61%, 73%, and 84% of the clip, so the samples span the full duration rather than clustering at the start.
  • Model: OpenAI CLIP ViT-B/32, INT8-quantized, in ONNX format. It maps images and text into the same 512-dimension space, which is what lets a typed phrase be compared against a frame at all.
  • Where it runs: on your Mac. Inference happens on-device through a statically linked ONNX runtime, so frames, vectors, and video files never leave the machine. There is a one-time download of about 149 MB for the model files themselves.
  • Vector storage: 16 KB per video (8 frames × 512 dimensions × 4 bytes). A 5,000-clip library adds about 80 MB of index.
  • How a search runs:your typed query goes through the same model's text encoder, then the app scans the vector table and scores each video by its closest-matching frame. On libraries of several hundred videos this returns results in well under a second.

Knowing that shapes what you type. The index holds a general impression of eight moments per clip, so a description of a whole scene lands, while a detail that flashed past for two seconds often falls between the samples.

What Content Search Finds – And What It Misses

Content search finds what a frame plainly looks like: a beach, a crowd, a laptop on a desk, a close-up of hands, an empty road. It misses what is not visible in the sampled frames – a brand name someone said out loud, text on a sign, or a two-second moment that fell between two of the eight samples.

Queries that work:

  • "aerial shot of a coastline"
  • "person walking a dog in the snow"
  • "coffee cup on a wooden table"
  • "night street with neon signs"

Queries that will disappoint you:

  • "the bit where Dan mentions the pricing" – that is speech, and there is no transcript.
  • "the slide that says Q3 Revenue" – that is on-screen text, and there is no character recognition.
  • "the client's logo in the corner" – small, low-contrast detail is not what a frame-level summary captures.
  • "the good take" – quality is a judgment call, and nothing in the frame encodes it.

Ranking is the other thing to expect. Results come back as an ordered list of what looks closest, so the clip you want usually sits near the top with its neighbors around it. For browsing footage that works in your favor, since you often want to see the alternatives anyway.

How To Search Your Library By Content On A Mac

Point a Mac app at the folders your footage already lives in, let it index in the background, then type what the shot looks like instead of what the file is called. In VidFinder that is three steps: connect your folders, wait for indexing, search a description.

VidFinder is a native macOS app that brings the videos from every hard drive, iCloud and Google Drive folder on your Mac into one visual library – scrub through previews – search by file name, folder name or simply describe what's in the shot. Find your clips in seconds, not minutes.

  1. 1

    Connect your folders. Add any folder Finder can see: a local folder, an iCloud or Google Drive folder, an external drive. Your files stay exactly where they are. VidFinder reads them and never moves, renames, or modifies an original.

  2. 2

    Let it index. The app generates a thumbnail and a short scrub-able preview for each clip, then builds the vectors described above. It reads 20 video container formats, including mp4, mov, mkv, mxf, mts, and avi.

  3. 3

    Search a description. Type what the shot looks like. Scrub the results by hovering a thumbnail, then drag the clip you want straight into your editor.

Because the previews and the index live locally, the library stays searchable when the drive holding the originals is unplugged. You can find and queue clips on a plane, then plug the drive in for the final drag.

When Filenames And Folders Still Win

Filenames still win when the name carries information a camera did not invent: a client name, a project code, a date you typed yourself, or a folder you created. Those names are exact, and an exact match beats an approximate one. Running both kinds of search in the same query is the setup worth having.

In VidFinder, keyword matching runs on every search regardless of the visual index. Each word you type is checked against the filename, your notes, your tags, and the full folder path. Multiple words narrow the results, so lisbon drone returns fewer clips than lisbon alone.

In practice, use the words you know are literally in the path or the tags to cut the library down, then use description to steer what surfaces inside it.


Frequently Asked Questions

Can Spotlight search inside videos on a Mac?

No. Spotlight indexes a video's filename and its file metadata – duration, codec, resolution, creation date – but not the visual content of the frames. Searching Spotlight for something you saw in a shot returns nothing unless that word happens to appear in the file's name or path.

Does searching video by content mean uploading my files somewhere?

It depends on the tool. Cloud services upload and process your footage on their servers. VidFinder runs the model on your Mac: frames, vectors, and originals never leave the machine, and the only network traffic is account sign-in and the one-time model download.

How much disk space does a content index take up?

In VidFinder, about 16 KB of vectors per video – roughly 16 MB per 1,000 clips – plus a one-time model download of about 149 MB. Preview clips and thumbnails are stored separately and are also small, since previews are 8-second, low-resolution files.

Can I search for words spoken in my videos?

Not in VidFinder. It indexes the visual layer only, with no speech transcription and no on-screen text recognition. For dialogue-heavy footage where you need a quoted line, a transcript-based tool is the right choice.

Will content search work if my external drive is unplugged?

Yes, for finding clips. The index and the previews are stored locally, so you can search, scrub, and select while the drive is disconnected. You need the drive connected to move or open the original file.

Do I have to tag my footage first?

No. The visual index is built automatically during indexing, so an untagged, camera-named library is searchable from the start. Tags are useful for the things a frame cannot show – client, project, usage rights – and those get picked up by keyword matching.

What video formats can be indexed?

VidFinder reads 20 container formats: mp4, mov, m4v, mkv, avi, webm, mts, m2ts, mxf, ts, 3gp, asf, dv, flv, gif, m2v, mpg, ogv, vob, and wmv. Camera-card structures and Final Cut, Photos, or iMovie libraries are treated as opaque – copy clips out of them to index them.

Search Your Own Footage By What's In It

Stop naming files you'll never search for. Point VidFinder at the drives you already have, and find the harbor shot by asking for the harbor shot.

Search My Own Library Free For 7 Days

7 days free · No credit card required · Cancel anytime