A duplicate file finder on a Mac: what to check before deleting
A duplicate file finder produces a number within a few minutes. Four thousand duplicates, twenty eight gigabytes recoverable, and a button that selects most of them. The number arrives with no explanation of how the tool decided, which volumes it looked at, or how much space those deletions would actually return. Treating that list as a verdict is where the damage happens. Treating it as evidence, and knowing what kind of evidence it is, turns the same list into a half hour of work with a predictable result.
What the tool is claiming when it says duplicate
Every duplicate finder implements one or more of four comparisons, and each answers a different question. Knowing which one produced a row explains both the false positives and the misses.
| Comparison | Catches | Gets wrong |
|---|---|---|
| Filename | Copies scattered across folders | Thousands of legitimate index.html and IMG_0001.JPG files in unrelated projects |
| Name and size | Whole folders that were duplicated | Misses a renamed report (1).pdf |
| Content hash | Real duplicates regardless of name | Treats a re-exported or re-compressed file as unrelated |
| Perceptual, for images and audio | Photos re-imported at another resolution | Flags crops and burst frames that were kept deliberately |
Filename matching is the fastest and the least trustworthy. Point it at a folder of source code and it returns hundreds of accurate, useless rows, because build systems and frameworks are designed around identically named files in different directories.
Content hashing is the only comparison that supports confident deletion. Two files with the same hash are interchangeable. Its weakness is that it is blind to intent: a photo saved again at lower quality, a video re-encoded for upload, a document exported to PDF twice on different days. A person sees one thing; the hash sees two unrelated files.
Perceptual matching has the opposite profile. It is the only way to catch the same image held at two resolutions, and its results cannot be verified from a text list. Whether two frames a second apart were kept on purpose is not answerable without looking at them.
Tools state this in their feature lists if you read for it. dupeGuru, distributed for Linux, Windows, and macOS 10.12 or later, documents that it scans either filenames or contents, that the filename scan uses fuzzy matching, and that it ships separate Music and Picture modes for tag and image comparison. That is three of the four comparisons in one application, which means the mode selected at scan time changes what the results mean.
The scan scope decides the answer
The second question to ask of a result list is what was in range. A duplicate is a relationship between two paths, so changing the set of paths changes the answer, and neither answer is wrong.
External drives are the clearest case. Scan the internal disk alone and an archived file reports as unique. Attach the archive drive and rescan, and the pair reports as redundant. The fix is to decide in advance which volume is the archive and which is the working copy, then delete only from the working side. A tool that lets you mark a folder as reference only, and refuses to select anything inside it, enforces that decision instead of relying on memory.
Synced folders need the same treatment for a different reason. Removing a local file in a synced directory removes it everywhere on the account. The candidate list gives no hint that a row behaves this way, so the scope has to exclude it or the review has to account for it.
Then there are the directories where a correct match is still the wrong thing to delete. The Library folder inside the home directory holds application settings, keys, licences, and mail data, and large files in there rarely have self explanatory names. A .git directory is project history, not a cache. The Photos library is a package whose internal records track its contents, so deleting a duplicate image from inside it breaks the library rather than shrinking it. Time Machine backup volumes are duplicates by design.
Some scanners exclude these by default and some do not. That default is worth checking before the first scan rather than discovering it in the results.
Two matches that free no space at all
The recoverable figure a scanner reports is a count of bytes in matched files, not a forecast of free space. On a current Mac, two reasons break that assumption.
APFS supports cloning. Duplicating a file in Finder, or copying it with cp -c, creates a second name pointing at the same blocks, and only later changes get written for real. Clone an 800MB file and the disk gives up a few megabytes, not another 800. A content comparison flags the pair as a duplicate, correctly, and du reports 800MB for each side, also correctly. Delete one and the free space figure barely moves. Hard links, used by some backup schemes, behave the same way.
The other reason is that deleted files are not gone. Apple states the first half of it directly in its storage guidance: moving a file to the Trash does not release its storage until the Trash is emptied. Local Time Machine snapshots extend the same effect for longer, which is what the purgeable portion of the free space readout refers to.
The practical habit is to stop counting rows. Note the figure in System Settings under General and then Storage, delete one batch, empty the Trash, and look again. The difference between those two readings is the only measurement of whether the approach is working.
Automatic selection, and the rule it is using
Most finders offer to select the copies to delete: keep the newest, keep the shortest name, keep the shallowest path, keep the one in a chosen folder. Each rule works exactly as described, and two of them tend to produce the reverse of what was wanted.
Date is the trap. Copying a file with cp writes the current time onto the destination while the original keeps its own timestamp. A document last edited in March and copied today produces a copy dated today and an original dated March. A rule that keeps the newer file keeps the copy. Files moved with cp -p or rsync -a preserve their timestamps, so one candidate list can hold files whose dates mean two different things.
Path depth is the other. Keep the shallowest and a re-downloaded file sitting loose in Downloads survives while the copy filed inside the project gets deleted.
A better rule uses references. The copy that something else points at is the one to keep: inside a project directory, inside a repository, inside a folder shared with other people. A reference names one path, so every other copy is unreferenced by definition. Downloads and the desktop are inboxes, and a file that has left the inbox and reached its place is the copy where the filing work has already been done.
Conflicted copies are the exception to every rule above. A file whose name contains a sync conflict marker holds edits that exist nowhere else. Those groups get opened and compared by hand, not selected by a rule.
The safeguards that matter in a tool
Feature lists are long, so it helps to know which four items change the outcome.
The first is a reference or protected folder, described above, which makes a scope decision enforceable.
The second is grouping that shows which file each match is being compared against, rather than a flat list of paths. dupeGuru describes both its reference directory system and its grouping as safety measures intended to prevent deleting files you did not mean to delete, which is a reasonable summary of why the feature exists at all.
The third is what the delete button does. Moving matches to the Trash leaves a recovery path; removing them outright does not. A related option in some tools is replacing a duplicate with a hard link or an alias, which keeps both paths working while storing the data once. That suits media libraries referenced from two places and is a poor fit for anything that will be edited later.
The fourth is the ability to export the candidate list. A list saved as text can be sorted by parent directory, and that single operation collapses most of the work. Two hundred scattered rows are hard to judge; the same rows re-sorted to show that one hundred and eighty of them live under a single old backup folder is one decision instead of two hundred.
Answering the same question with no installation
Running the built in tools first is worth it, because the result tells you whether a dedicated application is solving a problem you have.
find . -type f -size +1M -exec shasum -a 256 {} + | sort > /tmp/hashes.txt
awk '{ if ($1 == prev) print; prev = $1 }' /tmp/hashes.txt
du -sh */ | sort -h
find ~ -type f -size +500M
The first two lines hash everything above a size floor and print the groups that repeat. The floor matters twice over: hashing small files is slow, and matches among them free nothing. The last two answer the other question, which is where the space went rather than how many copies exist, and the storage breakdown in System Settings usually already indicates which of the two questions applies.
With Homebrew available, jdupes and rmlint add exclusion rules and hard link replacement, which suits a task that repeats monthly. Automation belongs at the end of this sequence, once the rule for choosing has been settled. Automating an unsettled rule deletes the wrong files faster.
Whichever route is used, do not chain a search into a delete. Write candidates to a file, read it, then act, or move them to a dated quarantine folder and revisit in two weeks. Removing a file with rm bypasses the Trash and cannot be undone.
Where the time actually goes
The detection is not the slow part. The round trip is. A group of candidates appears in a terminal, confirming what those files are means switching to a Finder window and pasting a path, and then the next group needs the same two moves. Each trip costs a few seconds, and three hundred candidates makes that an afternoon. It costs the same going the other way, since a large file spotted in a folder has to have its path copied before any command can touch it.
That is the case for a file manager with a built in terminal. When the listing and the prompt share one working directory, the path copy and the window switch disappear, and a candidate can be inspected, quarantined, and left behind in one place. The Features page sets out that layout, and Compared with other file managers places it against dual pane and terminal centric tools that address a different part of the job.
What to change first
Before reviewing any list, check three things about the scan that produced it: which comparison it used, which volumes were in range, and what the delete button does. Then measure free space before and after one batch rather than counting rows. If the friction that remains is switching windows to confirm each candidate, that is a layout problem rather than a detection problem, and Atriens is built around that gap.
Frequently asked questions
Why does a duplicate file finder report thousands of matches in a project folder?
Almost always because the scan ran in filename mode. Build systems and frameworks rely on identically named files in different directories, so index.html, main.js, and package.json appear in dozens of places legitimately. Switching the scan to content comparison removes those rows, and excluding node_modules and .git removes most of the rest.
Why did the recoverable figure not match the space actually freed?
The figure counts bytes in matched files. Space is not released until the Trash is emptied, and local snapshots can hold deleted data for longer. If the matches were APFS clones or hard links, two names pointed at one set of blocks and only one copy ever existed. Compare the storage figure before and after a batch to see which applies.
Is the automatic selection safe to use?
It is safe once you know which rule it applies. Keeping the newest file is the risky default, because a plain copy receives the current timestamp while the original keeps an older one, which makes the newer file the backup more often than not. Selecting by reference folder, where the archive side is protected, produces a more predictable result.
Do any of these tools handle the Photos library?
Most exclude it, and that is the correct behavior. The library is a package with its own records of what it holds, so removing an image from inside it leaves those records pointing at nothing. Duplicates inside the library are handled by the Duplicates collection in Photos, which merges rather than deletes and keeps the removed items recoverable for 30 days.