Commit Graph
85 Commits
Author SHA1 Message Date
TapTap 34970b961c feat: per-attribute preservation flags -p/-t/-o/-g with --no-* negations (protocol 2.22.0)
Split FastSync's single use_metadata bundle into four independent rsync-parity attributes: preserve_perms, preserve_times, preserve_owner, preserve_group. use_metadata is now a derived transport bit (config_derived_use_metadata).

CLI: real -p/--perms, -t/--times, -o/--owner, -g/--group plus --no-perms/--no-times/--no-owner/--no-group (short and long) and --no-preserve; -a is now rsync -rlptgoD; --preserve = -pt; -A implies -p; -X does not; --chmod implies -p; --usermap/--groupmap/--chown imply owner/group per side; --incremental/--delta still auto-preserve unless negated.

Receiver: per-attribute FileAttrPolicy gating for files, dirs (modes applied at end of transfer), symlinks and specials; rsync -E read-bit rule; new files get source_mode & ~umask sanitized (no group/other write); per-side identity resolution; deferred directory metadata; batch dir-metadata replay; daemon modules without 'client owner = yes' no longer refuse plain -a but force super off (no ownership) with a warning.

Wire: PROTOCOL_VERSION 2.21.0 -> 2.22.0 (four appended config bools, golden 653 / 95530566005420798). FileMetadata/chunk/batch framing unchanged. Docs/CHANGELOG/CMake updated to 2.22.0.
2026-09-15 19:32:02 +02:00
TapTap 5893de4a34 fix(receiver): close re-review findings — dry-run basis oracle, ACL capture, fsync reopen
Follow-up to a237043 addressing three security/correctness re-review findings.

(1) MEDIUM: a server-contacting --dry-run with --compare-dest/--copy-dest/
    --link-dest still read and hashed the basis file and compared it with the
    client-supplied digest, a 1-bit content oracle. basis_match_find() gains a
    hash_content parameter; the dry-run shortcut passes false and returns no
    match without touching basis bytes, so an otherwise-matching entry is
    reported as would-transfer. The real (non-dry-run) path is unchanged.

(2) LOW: xattr_capture_path() hardcoded preserve_acls=true, so the receiver's
    hard-link copy fallback re-applied system.posix_acl_* even when -A was not
    negotiated. The function now takes preserve_acls and members.* is
    unaffected; scanner and receiver callers thread the negotiated flag.

(3) INFO: the --fsync --link-dest temp reopen now uses O_NONBLOCK and treats
    a raced-in FIFO's ENXIO as a benign fsync-skip instead of blocking.

Tests: dry-run + basis unit test (asserts would-transfer, no content read) and
integration test; xattr capture ACL-filter test. Verified strict build, ASan,
clang-format, cppcheck, and the CI integration subset.
2026-09-14 17:19:27 +02:00
TapTap a2370433b2 fix(receiver): non-blocking receiver opens, inplace type gate, dry-run/B4/B5/B6
Address confirmed receiver security findings B1-B6:

B1 (HIGH): add O_NONBLOCK to the three receiver read-opens that opened an
existing destination/basis entry before the S_ISREG gate
(incremental_check_open_destination, basis_open_regular, hardlink_read_source)
so a client-planted FIFO can no longer block the receive thread forever while
the post-open type gate still rejects it.

B2 (HIGH/MED): --inplace now fstatat(AT_SYMLINK_NOFOLLOW)-probes the target and
refuses any existing non-regular entry, opens with O_NONBLOCK, and re-checks
S_ISREG on the opened fd.  This stops a FIFO from hanging the open and stops a
char/block device from being written directly (bypassing --write-devices).

B3 (MED): under --dry-run the incremental quick-skip no longer reads/hashes the
destination file for --checksum/--delta; it decides from metadata only and
reports would-transfer when the comparison is inconclusive, closing the
read-only-module content-hash oracle.

B4 (LOW): xattr_name_appliable() now gates the two system.posix_acl_* names on
preserve_acls (--acls), not the derived use_xattrs (--xattrs OR --acls).  The
receiver drops (never applies) ACL entries when -A was not negotiated while
keeping user.* working for -X.

B5 (INFO): receive_manifest_section() charges a per-entry overhead against
MAX_MANIFEST_BYTES and the aggregate entry count across all three sections is
capped at MAX_MANIFEST_ENTRIES.

B6 (MED): data_charge_session() reserves decompressed/chunk-copy bytes against
the owning ProtocolSession (MAX_CONNECTION_MEMORY) and records them on the Data
so data_destroy() releases them via the Data.owner path.  Applied to the
whole-file/append/delta decompression sites and chunk_deserialize() per-file
copies; a missing session owner degrades to the previous uncharged behavior.

Tests: FIFO destination/basis non-hang (with alarm), --inplace FIFO/device
refusal, dry-run no-read oracle test plus updated metadata-only dry-run tests,
ACL-without--acls drop, manifest total-entry cap, and chunk session charging.
2026-09-14 16:19:26 +02:00
TapTap 5d39619a8a Merge branch 'feat/w9-dryrun' into fix/w9-integration 2026-09-13 13:27:32 +02:00
TapTap 5b8aca5799 fix(receive): enforce dry-run no-mutation centrally
--dry-run --read-batch=FILE still wrote to the destination because
batch_read_apply -> file_save_to_disk_full bypassed the per-caller
!dry_run guards.  Guard file_save_to_disk_full and manifest_delete_all
directly (return SKIPPED/no-op) so every save/delete path is mutation-free
in dry-run, and keep the per-caller guards.  Reject --dry-run combined with
--read-batch/--only-write-batch at CLI validation with a clear error (a
dry-run of a local batch apply is not meaningful).
2026-09-13 12:57:30 +02:00
TapTap 88aee6ce94 feat(protocol): add optional STATUS_ERROR_DETAIL rejection reason (2.21.0)
Today a server rejection sends a bare STATUS_ERROR and the reason only
reaches the server log, so the client cannot say why a transfer was
refused.  Add an optional, bounded server->client error-detail frame:

  - Status gains STATUS_ERROR_DETAIL appended LAST so existing wire
    values are unchanged.
  - send_error_detail(fd, msg) sends STATUS_ERROR_DETAIL followed by the
    existing length-prefixed string primitive, slicing over-long messages
    to MAX_ERROR_DETAIL_BYTES (4096).
  - receive_status() (and the timed/keepalive status readers) always
    consume the detail body and map the status back to STATUS_ERROR,
    capturing the text into a thread-local buffer exposed by
    protocol_last_error(); a bare STATUS_ERROR leaves it cleared.  Every
    existing call site keeps working and the stream cannot desync.
  - Upgrade the daemon module gate / config validation (config.c), the
    final transfer failure (server.c) and receiver-side path/node
    validation (file_receive.c) to send a concrete reason; surface it on
    the client in client_send.c/config.c.
  - Bump PROTOCOL_VERSION to 2.21.0 (CMake VERSION, CHANGELOG, docs) and
    update the pinned config wire golden hash / CLI-version tests.
  - Add tests/test_protocol_error.c covering mapping+capture, the
    over-long bound, bare-error clearing, and thread-locality.
2026-09-13 12:19:46 +02:00
TapTap f6e8b6ddc4 refactor(file_receive): split receive_incremental_check into helpers
The per-file STATUS_CHECK fast path was a single 534-line function that
was hard to review.  Extract it into small static helpers called in order
by a short linear orchestrator:

  - incremental_check_receive_request  (receive/validate request frame)
  - incremental_check_open_destination (secure open + stat)
  - incremental_check_quick_skip       (metadata/content skip decision)
  - incremental_check_try_basis        (compare/copy/link-dest)
  - incremental_check_try_append_resume(--append tail resume)
  - incremental_check_try_delta        (block delta)
  - incremental_check_try_fuzzy        (--fuzzy basis)
  - incremental_check_receive_full     (STATUS_NEXT + whole file)

Pure refactor: the ordered sequence of wire operations
(send_status/send_n_data/receive_n_data/receive_status/receive_wire_str)
is byte-for-byte identical to the original, and every resource cleanup
is preserved (a unified idempotent cleanup replaces the duplicated
per-path close/free blocks).  No functional changes.
2026-09-13 12:04:31 +02:00
TapTap 6f974eff19 feat(dry-run): server-contacting --dry-run (protocol 2.21.0)
--dry-run now handshakes with a remote/daemon receiver and reports what
WOULD transfer/skip based on receiver state, mutating nothing on either
side.

- Serialize Config.dry_run into the wire config frame and append
  STATUS_DRY_RUN_TRANSFER to the status enum (no renumbering); bump
  PROTOCOL_VERSION/CMake VERSION/CHANGELOG/golden wire to 2.21.0.
- Receiver: receive_incremental_check_ex runs the normal read-only
  decision and answers STATUS_OK (skip) or STATUS_DRY_RUN_TRANSFER
  (would transfer) with no basis materialization/append/delta/full
  transfer.  All mutation sites are guarded by !dry_run: file store,
  manifest deletes, --mkpath root creation, --delay-updates staging,
  publication, directory-time application, and outcome acks.
- Client: send_dry_run_remote connects, sends the config, checks each
  regular file and prints the would-transfer set + trailer; no file data
  or delete manifest is sent.  Plain local destinations keep the
  client-side manifest.
2026-09-13 11:56:05 +02:00
TapTap ea2f76cd7a fix(receiver): charge per-entry DirTimeList cost; cap client --skip-compress 2026-09-13 01:18:54 +02:00
TapTap 4557924972 fix(receiver): cap DirTimeList growth and fix placeholder Data leaks 2026-09-13 00:57:32 +02:00
TapTap 4ac37c4d8a refactor(dir-times): extract dir_times_should_capture predicate
Deduplicate the repeated directory-time capture gate
(`config->use_metadata && !config->omit_dir_times`) used by the
sender-side (multiprocessing.c) and receiver-side (receiver.c) sinks
into a single predicate declared next to the DirTimeList machinery in
file_receive.h and defined in file_receive.c.

Behavior preserved: identical short-circuit condition and semantics,
no signature or protocol changes.
2026-09-12 20:42:47 +02:00
TapTap 6d32bc795b fix(p8h-core): escape log paths, fail closed on identity activation, tidy server gate
- A6: escape attacker-controlled file paths and the receive root in log
  lines (file_receive, server, protocol DEBUG) with output_escape()
- A8: identity_set_active() returns bool and fails closed when a requested
  usermap/groupmap cannot be deep-copied; handler refuses the connection
- remove the const cast and duplicate super_mode clamp from
  server_module_gate via an explicit override the handler applies once
- release the identity snapshot on the queue_create failure path
- refactor identity_parse_copy_as to a single cleanup tail and drop the
  duplicated group error format specifier
2026-09-12 15:03:54 +02:00
TapTap ea0a0e2eaf fix(p8-security): make --copy-as directory ownership airtight; harden tests/logs
- file_ensure_directory_secure() now chowns a final directory it creates under
  --copy-as and fails on error; the symlink parent-creation call site propagates
  it.  The is_dir branch fails when the confined parent cannot be opened under
  --copy-as.  Closes the residual wrong-owner gap for synthesized/symlink
  parent directories.
- file_restore_symlink_metadata() early NULL return is copy-as-aware.
- Preserve errno across the implicit-parent failure cleanup.
- Neutral skip messages (the clamp, not --no-super, may be responsible).
- Daemon copy-as test tolerates the non-root privilege refusal; usage text lists
  --copy-as.
2026-09-12 14:33:42 +02:00
TapTap b216ed31fb fix(p8-security): close review gaps in the ownership gate and copy-as failure propagation
- H3: a daemon module without 'client owner = yes' now also has super-user
  device activity forced off (char/block mknod, --write-devices), so a root
  daemon can no longer be made to create/write raw devices under AUTO.  The
  entries are skipped, preserving ordinary -a pushes.
- H1/H2: propagate a failed required --copy-as chown from symlink metadata
  restore and implicitly-created parent directories, so the entry (and run)
  reports failure instead of a wrong-owner success.
- Docs/help/headers updated for A2/A3 and the device clamp; startup warning
  spells out the client-owner risk.
- Tests: daemon device clamp (skipped without opt-in, created with opt-in),
  updated --super/--fake-super expectations.
2026-09-12 14:18:03 +02:00
TapTap e3840c8326 fix(p8-security): enforce daemon ownership policy, gate fake-super replay, drop implicit numeric-ids, make copy-as failures per-entry
A1: daemon refuses every client-chosen ownership/super-user request
(--numeric-ids/--chown/--usermap/--groupmap/--fake-super/--copy-as/--super)
unless the selected module opts in with 'client owner = yes'.
A2: fake-super owner replay requires an explicit ownership identity policy.
A3: --super no longer implies --numeric-ids (ownership stays opt-in).
A5: a failed --copy-as chown marks the entry failed instead of reporting
success with the wrong owner.
2026-09-12 14:00:55 +02:00
TapTap 938886829e fix(p7-privilege): close re-review gaps (implicit dir ownership, daemon --super, write-devices gate) 2026-09-12 12:54:45 +02:00
TapTap fdc0f238c9 fix(p7-privilege): harden copy-as/super gates, own dirs/specials
- fake-super owner replay honors --no-super and an active --copy-as
- copy-as/identity ownership now applied to directories and special nodes
- reject copy_as_set && !use_metadata (receiver + client --no-preserve)
- daemon refuses --copy-as; add server-side --no-super operator veto
- implement identity_copy_as_refused/identity_copy_as_active
- reject copy-as ids that overflow int32; escape spec in log errors
- copy-as chown EPERM/EACCES logged at ERROR (still non-fatal)
- identity_wire_valid copy-as bounds; CLI help and RSYNC_COMPAT docs
- add unit tests and root-gated integration coverage
2026-09-12 12:34:43 +02:00
TapTap a785ec13c4 feat(p7-super): implement --super/--no-super safe-subset privilege gate (protocol 2.18.0)
Add the receiver-side --super / --no-super tri-state (Config->super_mode)
under the safe-subset + clear-refusal privilege model: FastSync never
elevates privileges, it only permits super-user attempts that are already
confined fd-relative below the authorized receive root.

- identity: privilege_super_permitted() gate (OFF=false, ON=true, AUTO follows
  geteuid()==0); identity_apply_ownership/_link become no-ops when not
  permitted; --super with no explicit identity policy implies raw numeric-id
  preservation (explicit usermap/groupmap/chown/numeric-ids still win); warn
  exactly once when --super is requested by a non-root receiver.
- file_receive: gate char/block device-node creation on the gate; FIFO/socket
  handling is unchanged.
- wire: trailing super_mode int after the --iconv spec, validated 0..2 in
  receive_privilege_options and validate_received_config; PROTOCOL_VERSION
  2.17.0 -> 2.18.0; version-sensitive tests and docs updated.
- CLI: --super/--no-super parsed explicitly before the generic --no-* branch
  (malformed --super=x rejected); usage text added.
- tests: config wire round-trip + invalid-value rejection, privilege-gate mode
  unit test, CLI parse test, integration transfer + root-gated ownership
  suppression/appliance tests.
- docs: RSYNC_COMPAT --super row + Wave E note, protocol mentions, README.
2026-09-12 12:01:19 +02:00
TapTap 003a5e8f2f Merge feat/p7-times: Phase 7 Wave D (real dir/symlink time preservation making -O/-J meaningful; secluded-args -> Impossible/Divergence; PROTOCOL 2.17.0) 2026-09-12 11:22:23 +02:00
TapTap 9749c7878c fix(p7-times): dir-time entries only record (never create dirs); bound/chunk dir-time frames; harden list add; docs+tests
Review fixes for Phase 7 Wave D.

#1 (HIGH): STATUS_DIR_TIMES entries no longer create directories. A new
receiver-only File.dir_time_only flag marks dir-time entries; file_save_to_disk_full
short-circuits them as FILE_SAVE_SKIPPED before any device/dir branch, so the sink
still accumulates metadata into the deferred DirTimeList but creates nothing. Empty
source dirs stay untransferred (-a), -m/--prune-empty-dirs semantics are preserved,
and a pre-existing regular file/symlink at an empty-dir mirror path no longer aborts
the transfer. dir_time_list_apply fstatat()s the leaf (AT_SYMLINK_NOFOLLOW) and skips
absent/non-directory paths QUIETLY; only a real existing directory is stamped.
Also initialize File.dir_time_only in file_create() (uninitialised garbage otherwise).

#2 (MED): send_dir_times() chunks entries into repeated STATUS_DIR_TIMES frames of at
most MAX_MANIFEST_ENTRIES, matching the receiver's per-frame bound; the tautological
> INT_MAX check is gone.

#3 (LOW): dir_time_list_add() assigns each grown array right after its realloc (no
dangling) and advances capacity only after both succeed.

#4 (LOW): RSYNC_COMPAT.md -- STATUS_MKDIR carries metadata, dir times are transmitted
via STATUS_DIR_TIMES and applied at the end, empty dirs are still never created; -m
rationale, -O row and Wave D notes updated. Summary counts untouched.

#5 (LOW): integration tests for the three #1 scenarios (empty-dir non-creation under
-a and -a -m, collision non-abort), scanner test now covers empty-dir capture, and
test_file_restore_symlink_metadata asserts the positive apply path when supported.

PROTOCOL_VERSION stays 2.17.0; config-frame layout unchanged.
2026-09-12 11:06:05 +02:00
TapTap f1a447bb4a feat(p7-times): real directory/symlink time preservation; -O/-J meaningful
Wave D of Phase 7. Make -O/--omit-dir-times and -J/--omit-link-times real by
preserving directory and symlink times, and mark --secluded-args as an explicit
Impossible/Divergence no-op.

Wire: PROTOCOL_VERSION 2.16.0 -> 2.17.0. Adds a terminal STATUS_DIR_TIMES frame
(int count + (wire path, metadata) pairs) sent after all file data and the
optional delete manifest. STATUS_MKDIR also carries metadata for --dirs entries.
Config-frame layout is unchanged.

Sender: the recursive scanner captures every traversed source directory (both
DirectoryScanner and the parallel scanner root + workers, appends mutex-guarded)
into a shared list; the single-threaded and -m paths transmit it last.

Receiver: a DirTimeList accumulates received directory metadata and applies it
with fd-relative no-follow utimensat only at the very end -- after all children,
after the commit-style --delete, and after --delay-updates publication -- in the
single-threaded success frame and in server.c after the -m threads join. -O skips
the application. Symlink metadata is applied at link creation with
utimensat/fchownat/fchmodat AT_SYMLINK_NOFOLLOW; -J suppresses only link times.
identity_apply_ownership_link shares the identity resolver with the fd path.

Docs: -O/-J rows -> Implemented; --secluded-args -> Impossible/Divergence;
--protocol accepted/rejected values and Phase-6/7 notes updated.

Tests: unit (scanner dir capture, DirTimeList apply, symlink metadata, protocol
version values) and integration (dir mtime round-trip + -O, symlink mtime
round-trip + -J, independent suppression), parameterized over single/multithread.
2026-09-12 10:31:20 +02:00
TapTap 47de05d215 feat(p7-output-fs): sparse hole preservation (-S), partial retention (-P), fake-super replay; block-size verified; crtimes/stderr -> Impossible/Divergence (Wave B)
- write_all_sparse: skips all-zero runs >= 4096 bytes via lseek(SEEK_CUR) and
  ftruncates the final size, wired into the atomic temp+rename and --inplace
  paths with no wire change (full image already in memory).
- --partial retention: on a save failure after the temp held data, rename the
  already-written temp to the destination path (best-effort; falls through to
  unlink; never retains when --partial is off) so --append/--append-verify can
  resume; tested by forcing futimens EINVAL with an out-of-range nsec.
- --block-size aliases --delta-block; verified config->delta_block_size is
  honored by the delta engine end-to-end (unit + integration tests).
- fake_super_restore_fd: parses and re-applies user.fastsync.stat fd-relative
  (fchown best-effort/non-root skipped, fchmod, futimens); a save under
  --fake-super now both records and re-applies.
- -N/--crtimes and --stderr=client promoted to a new 'Impossible/Divergence'
  status bucket (Summary: 136/2/4/3/2 = 147).
- cppcheck/clang-format clean; unit 37/37; integration 407 passed.
2026-09-11 15:58:20 +02:00
TapTap 6e02a24232 feat(p6-iconv): --iconv charset conversion + PROTOCOL 2.16.0 2026-09-10 17:13:01 +02:00
TapTap ad228db915 fix(p5-remote-option): wire server --trust-sender, align save-layer gates, add hostile-sender test 2026-09-09 14:23:59 +02:00
TapTap 5e79d7d76b feat(p5-remote-option): --remote-option (probe 2.14.0), --trust-sender 2026-09-09 13:41:28 +02:00
TapTap 94b6662018 Merge feat/p4-acl-xattr: -X/-A/--fake-super
# Conflicts:
#	src/shared/config.c
#	src/shared/file.c
#	src/shared/file_types.h
#	tests/integration/test_features.py
#	tests/test_client_cli.c
#	tests/test_config.c
2026-09-08 22:38:15 +02:00
TapTap 5bbdc5b450 Merge feat/p4-devices
# Conflicts:
#	src/client/client_send.c
#	src/server/receiver.c
#	src/shared/chunk.c
#	src/shared/file.c
#	src/shared/file_receive.c
#	src/shared/file_receive.h
#	src/shared/file_types.h
#	src/shared/protocol.h
#	tests/test_chunk.c
2026-09-08 22:33:52 +02:00
TapTap 747946c318 acls/xattrs: -X/--xattrs, -A/--acls, --fake-super
CI / lint (pull_request) Failing after 55s
CI / build-and-test (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
New src/shared/xattr.{c,h}: capture user.* + POSIX ACL xattrs, transmit a bounded
per-file block, re-apply fd-relative. security.*/trusted.*/other system.* never
transmitted/applied (receiver re-validates). Bounds: name<=255 value<=1MiB count
<=256 total<=4MiB. --fake-super records uid:gid:mode:mtime in reserved
user.fastsync.stat (receiver-only). PROTOCOL_VERSION 2.12.0->2.13.0. Review
fixes: reserved key not forwardable, link/hardlink copy-fallback preserves
xattrs, no const-param mutation, per-file warning dedup.
2026-09-08 22:27:53 +02:00
TapTap 007e8f90f2 devices: --devices/--specials/-D/--copy-devices/--write-devices
CI / lint (pull_request) Failing after 56s
CI / build-and-test (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
Recreate char/block nodes via mknodat (privilege-gated, EPERM->warn+skip) and
FIFOs via mkfifoat; new STATUS_SPECIAL frame + validated rdev; sockets skipped;
-copy-devices copies st_size; -write-devices O_NOFOLLOW+O_NONBLOCK warn+skip.
preserve_specials/copy_devices/write_devices cross the wire. PROTOCOL_VERSION
2.12.0->2.13.0. Review fixes: -m source-removal keeps recreated specials, FIFO
ENXIO skip, rdev bounds at chunk_deserialize, STATUS_ERROR on receive branch,
scanner_prepare_special dedup.
2026-09-08 22:27:53 +02:00
TapTap 820188c2cc symlink-trust: -k/--copy-dirlinks, -K/--keep-dirlinks, --munge-links (+real -l/--links)
CI / lint (pull_request) Successful in 1m1s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m19s
Adds symlink-target transmission (File is_symlink+symlink_target, STATUS_SYMLINK
frame, chunk type 2), munge-links sender containment + receiver-side symmetric
target containment, keep-dirlinks confined dir-symlink following (O_NOFOLLOW
realpath-rechecked), and fixes -l to copy symlinks as symlinks. munge_links +
keep_dirlinks cross the wire; copy_dirlinks client-only. PROTOCOL_VERSION
2.12.0->2.13.0. Review fixes: receiver rejects absolute/.. targets, gated unmunge,
-K O_NOFOLLOW+re-fstat, keep_dirlinks set once at config-accept, rel_buf overflow
fails the walk.
2026-09-08 22:27:53 +02:00
TapTap 4e7e84f947 Merge feat/p4-metadata-capture: atimes/crtimes/open-noatime/omit-dir-times/omit-link-times
# Conflicts:
#	src/shared/config.c
#	tests/integration/test_features.py
2026-09-08 21:00:07 +02:00
TapTap e80888ce7b metadata times: -U/--atimes, -N/--crtimes, --open-noatime, -O/-J
CI / lint (pull_request) Successful in 52s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m19s
Capture+transmit source atime (pre-read stat; O_NOATIME sender guard) and birth
time (statx STATX_BTIME); receiver restores atime with mtime (crtime not settable
portably -> transmitted, explicitly not applied). --open-noatime is client-only.
-O/-J documented as accepted no-ops (FastSync never preserves dir/symlink times).
Wire: metadata frame gains atime/crtime val+sec+nsec; PROTOCOL_VERSION
2.11.0->2.12.0. Review fixes: gate atime capture to Linux (no epoch clobber on
non-Linux), close fd on fdopen failure, honest -O/-J status (Compat no-op).
2026-09-08 20:58:56 +02:00
TapTap f891cd0a6a hard-links: -H/--hard-links preserves inode relationships
CI / lint (pull_request) Successful in 50s
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
CI / build-and-test (pull_request) Successful in 1m19s
Source files sharing (st_dev,st_ino) are recreated as hard links on the
destination; only the first member's data crosses the wire (siblings ride a
payload-less STATUS_HARDLINK frame). Ordering requires the single-FIFO-writer
receiver + forced sequential scan (documented). link()-failure falls back to a
byte-identical local copy. Rejects -s/--append. PROTOCOL_VERSION 2.11.0->2.12.0.
Review fixes: delete the dead HardLinkRegistry (ordering holds by FIFO writer),
and --existing no longer aborts when the first member is absent but the sibling
exists (leaves the sibling in place).
2026-09-08 20:58:56 +02:00
TapTap 362a6a5488 preallocate: --preallocate allocates dest space up front
CI / lint (pull_request) Failing after 3s
CI / build-and-test (pull_request) Skipped
CI / sanitizers (address) (pull_request) Skipped
CI / sanitizers (undefined) (pull_request) Skipped
CI / fuzz-build (pull_request) Skipped
CI / coverage (pull_request) Skipped
CI / valgrind (pull_request) Skipped
Receiver allocates the destination file's full size before streaming data
(posix_fallocate preferred, ftruncate fallback on EOPNOTSUPP/ENOSYS) so an
out-of-space transfer fails fast instead of partway. Additive config bool
crossing the wire; PROTOCOL_VERSION 2.10.0 -> 2.11.0. Threaded through all
store paths (atomic, inplace, partial/delay-updates staging, link-dest copy
fallback). Review hardening: explicit lseek(0) before the data write so
correctness does not depend on posix_fallocate leaving the fd offset unchanged.
2026-09-08 18:23:48 +02:00
TapTap 4c98744014 Merge feat/p3-checksum-choice: --checksum-choice/--cc and --checksum-seed 2026-09-07 19:53:28 +02:00
TapTap f99e6e1edf Merge feat/p3-append: --append / --append-verify tail resume 2026-09-07 19:52:33 +02:00
TapTap c6758531f6 fix: reject xxh3 alias, pin handshake digest length to algorithm, simplify xxh64 copy, note FIPS md5
CI / lint (pull_request) Successful in 35s
CI / sanitizers (undefined) (pull_request) Successful in 52s
CI / sanitizers (address) (pull_request) Successful in 52s
CI / fuzz-build (pull_request) Successful in 18s
CI / coverage (pull_request) Successful in 42s
CI / valgrind (pull_request) Successful in 37s
CI / build-and-test (pull_request) Successful in 12m58s
2026-09-07 19:50:52 +02:00
TapTap 4b74f6a41d fix: treat an absent missing-mirror parent as a no-op, only claim real deletions
BLOCKER: an absent mirror whose PARENT directory does not exist on the
destination (a deeper --files-from missing entry, -R or full-mirror layout) was
treated as a hard failure because file_open_secure_parent returned -1 when the
parent was missing. That aborted the whole run, skipped the --delete extras
walk, and tore down an early --delete-before/during connection, contradicting
'missing mirror = no-op / all-missing succeeds'. A parent-open failure is now a
no-op when errno is ENOENT/ENOTDIR (matching file_remove_tree_secure); only a
genuine I/O error fails the run.

Also: an already-absent mirror reached via unlinkat-ENOENT no longer prints
'Deleted: <path>' (a no-op dressed as a deletion); the per-path 'Deleted:' line
is printed only when an entry was actually removed.
2026-09-07 18:32:47 +02:00
TapTap 0afda6b094 feat(checksum): --checksum-choice/--cc and --checksum-seed for whole-file digest
Adds real algorithm selection (xxh64 default, plus md5 via OpenSSL EVP) and a
64-bit seed for the per-file whole-file digest used by the --incremental/
--checksum handshake and basis-dir content verification. The seed also feeds
the delta path's per-block xxHash32 strong checksum (low 32 bits) so an
explicit seed deterministically changes those digests too. Sender and receiver
hash identically: the algorithm id and seed cross the config wire frame and the
STATUS_CHECK handshake now carries a length-prefixed, bounded digest instead of
a fixed 64-bit value. Unsupported algorithm names are rejected at parse time
(never a silent no-op). PROTOCOL_VERSION bumped 2.9.0 -> 2.10.0; defaults
(xxh64, seed 0) preserve prior byte-for-byte behavior.
2026-09-07 17:42:47 +02:00
TapTap 6ad3887aa1 feat: --append / --append-verify tail-only resume
Implement rsync's append modes: when an existing destination file is SHORTER
than the source, the receiver negotiates a resume offset and only the tail is
transferred; the full file (retained prefix + tail) is rebuilt and installed
through the normal atomic store path, so the result is byte-identical to the
source whenever the prefix matches.

- --append: sends the tail without content-verifying the retained prefix
  (rsync parity; the documented prefix-trust risk).
- --append-verify: verifies the retained prefix against the source's prefix
  xxHash64 before appending and, on a mismatch, falls back to a clean full
  transfer (never a corrupt prefix+tail blend).

New wire frames STATUS_APPEND / STATUS_APPEND_SIG / STATUS_APPEND_OK /
STATUS_APPEND_DATA; PROTOCOL_VERSION bumped 2.9.0 -> 2.10.0 (peers must match).
Both flags imply --incremental and are incompatible with -s (chunk
serialization) and --whole-file (rejected up front). Respects --inplace,
--partial/--partial-dir and --delay-updates via the shared store engine.
2026-09-07 15:43:29 +02:00
TapTap c9bd76e633 feat(shared): --delete-missing-args config/wire, manifest third section, receiver exact-path deletions
PROTOCOL_VERSION 2.9.0 -> 2.10.0. The STATUS_MANIFEST frame gains a third
section carrying destination-relative exact-delete paths (the missing
--files-from entries' mirrors); the config frame gains a delete_missing_args
bool (ignore_missing_args stays client-only). The receiver validates the third
section like the keep-set and commits it with manifest_delete_all():
manifest_delete_missing_args runs first (explicit user requests, never blocked
by protected-prefix exclusion protection; staging/basis protected; a non-empty
directory mirror removed only under --force/--delete, rsync parity) and then the
ordinary extras walk. Server --allow-delete gates it like --delete.
2026-09-07 14:34:19 +02:00
TapTap ebab335488 Merge feat/p3-fuzzy: --fuzzy/-y/--no-fuzzy similar-file delta basis 2026-09-06 23:52:24 +02:00
TapTap 87b58975de style(receiver): rework fuzzy DP suffix trim, silence cppcheck FP
CI / lint (pull_request) Successful in 37s
CI / sanitizers (undefined) (pull_request) Successful in 42s
CI / sanitizers (address) (pull_request) Successful in 44s
CI / fuzz-build (pull_request) Successful in 17s
CI / coverage (pull_request) Successful in 34s
CI / valgrind (pull_request) Successful in 36s
CI / build-and-test (pull_request) Successful in 7m19s
The suffix-trim loop using computed end offsets tripped cppcheck's
knownConditionTrueFalse value-range analysis (it unsoundly concluded the trims
always consume the whole middle).  Rewrite it with explicit moving end indices
and add an inline suppression with a rationale for the residual false
positive; cppcheck --error-exitcode=1 is clean again.  The trimming logic is
unchanged and was verified against a full DP reference over 200k random name
pairs.
2026-09-06 23:11:01 +02:00
TapTap be37a509a5 perf(receiver): bound the fuzzy edit-distance cost, harden the open
Review follow-ups on the --fuzzy candidate scan:
- Allocate the two DP rows once per directory scan instead of once per
  candidate (4096-entry directories no longer do thousands of malloc pairs).
- Pre-prune before the DP with two cheap lower bounds on the edit distance:
  the name length gap and the count of characters of one basename absent from
  the other; a candidate whose gate (distance*2 <= longer) already fails on
  the max of those bounds is skipped without running the DP.
- Trim the common prefix and non-overlapping common suffix before the DP so
  it only runs over the differing middles.
- Note that the 4096 readdir cap bounds iterations, not per-entry DP cost,
  and that the seen set is filesystem-order dependent (winner stays
  deterministic via the total comparator).
- Open the chosen candidate with O_NONBLOCK so a name raced to a FIFO cannot
  block the receive thread forever in open(2); the existing fstat S_ISREG gate
  still rejects non-regular files.  (The pre-existing basis_open_regular has
  the same latent FIFO pattern and is intentionally left unchanged.)
- Free old_data in receive_delta_file's defensive NULL guard.
2026-09-06 23:05:32 +02:00
TapTap c4e0de8f08 feat: split delete-manifest frame into keep-set + protected prefixes; enforce --max-delete and --force on the receiver
The STATUS_MANIFEST frame now carries two count-delimited sections: the kept
paths and a protected-prefix list (excluded-on-source paths the walker must not
delete unless --delete-excluded opted out).  The receiver's DeleteManifest is
passed through the commit/early paths unchanged.  manifest_delete_extras
honors a client --max-delete (all-or-nothing) and produces a distinct error for
it versus the 100000-entry server bound.  --force clears a non-empty directory
that blocks an incoming regular file (confined, symlink-safe) via a new
file_remove_tree_secure helper.
2026-09-06 21:50:02 +02:00
TapTap cebd23a239 feat(receiver): --fuzzy similar-file delta basis
When a file must be transferred and the destination holds no usable content
at the exact path (file absent, or outside the delta engine's size bounds),
the receiver now searches the target's own destination directory for an
existing regular file with a similar basename and uses it as the delta basis
via the existing receiver-driven STATUS_DELTA_SIGNATURE handshake.  The
sender never learns the basis was another file, so no wire change beyond the
new config flag was required.

Heuristic (deterministic, simpler than rsync's, documented): candidates are
sibling entries confined below the root and opened O_NOFOLLOW (symlinks are
never followed, nothing outside the destination root is read); dotfiles,
directories, the target's own name and stage/temp names are excluded; size
gate is delta_should_attempt; name gate is a Levenshtein distance <= half
the longer basename; the closest candidate (size tie-break, then lexical) is
loaded; the scan is capped at 4096 entries.

Byte-exactness is independent of the basis: block matches are verified by
Adler-32 + xxHash32, delta_apply validates every reference, and a basis that
shares nothing makes the sender reply with a whole-file transfer.  When no
candidate qualifies the normal whole-file transfer runs unchanged.
2026-09-06 21:04:19 +02:00
TapTap 01a5a93089 Merge feat/p3-basis-dest: alternate basis dirs (--compare-dest/--copy-dest/--link-dest) 2026-09-06 20:10:40 +02:00
TapTap 4bf4da37e5 fix: free basis path on copy-dest hits; keep --delete staging skip top-level-only
- copy-dest basis hits leaked the heap-allocated basis path: BASIS_DEST_COPY
  did not transfer it (only LINK does) and returned before the basis cleanup.
  basis_match_free is now called on every materialization return path (success
  and send-failure) after content/link ownership is transferred.
- the --delete walker regression: the delay-updates staging name must be
  protected only as a DIRECT child of the receive root, while basis dirs may
  be skipped at any depth.  delete_extras_limited now takes DeleteSkipEntry
  entries carrying a top_level_only flag instead of a flat prefix list, so a
  nested destination directory named .fastsync-stage is ordinary content again
  (its extras are deleted) and a basis tree is still never removed.
2026-09-06 19:43:11 +02:00
TapTap 4e725517f0 feat: implement rsync delete timing (--delete-before/--delete-during/--delete-delay/--delete-after)
Deletion timing is now real and selected by the four rsync flags plus the
plain --delete default. Wire protocol bumps to 2.8.0: two new config
booleans (delete_during, delete_delay) are serialized and validated, joining
the existing delete_before/delete_after.

- Early modes (--delete-before, --delete-during/--del): the sender pre-scans
  the whole tree (paths only), transmits the keep-set manifest BEFORE any
  file data, and the receiver removes extras and acks STATUS_OK; the sender
  only streams data after the deletion committed. Deletion is thus performed
  even if a later transfer phase fails (rsync delete-before/during are
  destructive by definition). FastSync streams in a single scan so it cannot
  interleave per-directory like rsync delete-during; --delete-during selects
  the same engine mode as --delete-before (documented divergence).
- Late/commit modes (plain --delete, --delete-after, --delete-delay): the
  manifest closes the data stream and deletion is committed only after
  STATUS_FINISHED proves the whole transfer succeeded, preserving FastSync's
  commit-style safety. --delete-delay converges with --delete-after because
  FastSync never snapshots the destination during data flow (documented).
- The STATUS_MANIFEST frame is now self-delimiting and position-independent.
  Single-threaded receivers delete before the success frame; the -m receiver
  hands the keep-set to server.c, which commits the deletion only after the
  disk writer thread has drained (fixes a delete-vs-in-flight-temp race).
- Every timing flag implies --delete; at most one timing flag is allowed.
- Each timing flag implies --delete, matching rsync; conflicts are rejected.
2026-09-06 18:53:20 +02:00
TapTap d38920c972 feat: receiver-side basis matching with link/copy materialization
The receiver's per-file incremental check now consults the ordered basis-dir
list whenever the destination is not already up to date.  An exact basis match
requires equal size, equal mtime (unless --size-only; --ignore-times disables
basis matching like rsync), and an equal content xxHash64 -- the sender sends
its xxHash for every file whenever basis dirs are configured (not only under
--checksum), so a hard link or local copy is only ever made from byte-identical
content.

On a match:
  - compare-dest: reply STATUS_OK and skip data only when the destination does
    not already hold the file (sparse, rsync parity).  A destination that holds
    a DIFFERENT version falls back to a normal transfer instead of rsync's
    delete, keeping the mirror complete.
  - copy-dest: reply STATUS_OK and hand a synthetic File (bytes read from the
    basis file, basis metadata) to the normal store sink, so the file is
    installed as a real local copy through the existing atomic temp+rename
    engine and honors --existing/--ignore-existing/--update/--backup/
    --delay-updates/--partial-dir unchanged.
  - link-dest: same, but File.basis_link records the basis path and the store
    engine calls the new file_to_disk_secure_link(): an atomic temp hard link +
    rename.  Cross-filesystem/refused links fall back to a byte-identical local
    copy (never a corrupt or partial file); the copy fallback applies metadata,
    while a successful link keeps the basis inode's own attributes so the basis
    file is never mutated.

Basis-materialized files carry File.skip so they are not acknowledged to a
--remove-source-files sender (the sender already saw STATUS_OK and keeps the
source).  The no-match path is byte-for-byte identical to the existing delta /
full-data transfer.
2026-09-06 18:47:26 +02:00