Skip to content

Never treat one file reached at two paths as a duplicate of itself [patch] - #181

Merged
matt-edmondson merged 3 commits into
mainfrom
fix/same-file-at-two-paths
Oct 8, 2026
Merged

matt-edmondson merged 3 commits into
mainfrom
fix/same-file-at-two-paths

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #139
Fixes #132

What was wrong

A bind mount, a share mounted twice, or a hardlink gives one file several paths. None of these is a reparse point, so the scan listed the file once per path. The paths hashed the same, landed in one duplicate group, and Deduplicate deleted the "other copy". Through a bind mount, that copy is the keeper's own directory entry, so the only copy of the data was deleted and reported as reclaimed space (#139). Through a hardlink, the deleted path was a name in another snapshot, and the reported reclaimed bytes were wrong (#132).

What changed

  • FileIdentity (new) reads a file's (device, index) identity:
    • On Unix, (st_dev, st_ino) through the runtime's SystemNative_Stat shim. FileType uses the same shim for the same reason: its structure has fixed-width fields on every Unix.
    • On Windows, (VolumeSerialNumber, FileIndexHigh/Low) from GetFileInformationByHandle.
  • Deduplicator.FindDuplicates keeps one path per identity within each hash group, the path SelectFileToKeep would prefer. As a result, Scan, DryRun, Stats and Deduplicate no longer report one file seen at two paths as a duplicate group. Identity is read only for files already in a group of two or more, so the scan itself is not slowed.
  • Deduplicator.DeleteDuplicates gets the last guard the triage suggested: a candidate whose identity equals the keeper's is never deleted. It is reported as a SkippedFile instead, whatever grouping the loop is handed.
  • The README's "Which File Is Kept?" section explains the rule.

For hardlinks, this takes #132's first option: an extra hardlink is treated as the same file, matching the existing symlink policy.

Tests

New SameFileTests:

  • Hardlinks share an identity and same-content copies do not. On Unix, the index equals the inode number ls -i reports, which pins the struct offsets.
  • A missing file has no identity, and reading it does not throw.
  • A file and its hardlink form no duplicate group. Next to a genuine copy, the hardlink is collapsed, leaving a group of two.
  • DeleteDuplicates, given a hand-built group of a keeper and its hardlink, deletes nothing and reports the link as skipped.
  • The issue's reproduction, end to end: a real mount --bind of real/ onto view/, then Deduplicate confirmed with y. The file survives with its content, nothing is deleted, and DryRun reports no duplicates. The test is inconclusive where bind mounts are unavailable (non-Linux, or not root).

With the Deduplicator.cs change reverted, 4 tests fail: both grouping tests, the deletion-guard test and the bind-mount repro. The full suite passes locally on Linux: 97 passed, 3 inconclusive. Those 3 were already inconclusive because they need an unprivileged user. The Windows hardlink path (mklink /H) and the GetFileInformationByHandle path run only in CI.

Note

PR #180 changes TryGetSize's visibility and the verbs. This PR changes neither, so the two should merge cleanly in either order.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EaKE6BddCbJMXwfuLwNnGh


Generated by Claude Code

…atch]

A bind mount, a share mounted twice or a hardlink gives one file several
paths. They hashed the same and were grouped as duplicates, and Deduplicate
then deleted the "other copy". Through a bind mount that copy is the keeper's
own directory entry, so the only copy of the data was deleted and reported as
reclaimed space. The keeper re-hash could not catch it, because the keeper was
the file being deleted.

FileIdentity reads a file's (device, index) identity: (st_dev, st_ino)
through the runtime's SystemNative_Stat shim on Unix, and the volume serial
number and file index from GetFileInformationByHandle on Windows.
FindDuplicates keeps one path per identity, the one SelectFileToKeep would
prefer, so Scan, DryRun, Stats and Deduplicate no longer report such a file as
a duplicate group. As a last guard, the deletion loop skips any candidate whose
identity equals the keeper's and reports it as a skipped file.

Fixes #139
Fixes #132

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EaKE6BddCbJMXwfuLwNnGh
Comment thread FileDeduplicator/Deduplicator.cs Fixed
claude added 2 commits October 7, 2026 21:30
…ty under the limit

SonarCloud S3776 counted 18 against the allowed 15 once the identity guard was
added. The try/catch around File.Delete moves into TryDelete unchanged, and the
Unix-only identity test uses [OSCondition] instead of an early Inconclusive.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01EaKE6BddCbJMXwfuLwNnGh
@sonarqubecloud

sonarqubecloud Bot commented Oct 7, 2026

Copy link
Copy Markdown

@matt-edmondson
matt-edmondson merged commit 473e1d0 into main Oct 8, 2026
14 checks passed
@matt-edmondson
matt-edmondson deleted the fix/same-file-at-two-paths branch October 8, 2026 04:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants