receiver_save_file appended to context->outcomes for --remove-source-files
without the !dry_run guard the multithreaded pipeline has, so a hostile
dry-run client could grow outcomes unbounded (raw, uncharged realloc) and
force a per-frame ack. Guard the append on !dry_run.
--dry-run --read-batch=FILE still wrote to the destination because
batch_read_apply -> file_save_to_disk_full bypassed the per-caller
!dry_run guards. Guard file_save_to_disk_full and manifest_delete_all
directly (return SKIPPED/no-op) so every save/delete path is mutation-free
in dry-run, and keep the per-caller guards. Reject --dry-run combined with
--read-batch/--only-write-batch at CLI validation with a clear error (a
dry-run of a local batch apply is not meaningful).
--dry-run --server-host=H (or TLS / source-bind --address) silently ran the
client-side manifest even though a real run contacts the server. Add a
client-only, never-serialized server_host_set bit (alongside the existing
server_port_set) and extend dry_run_targets_server so every explicit remote
target contacts the receiver.
Also make incremental_check return the dry-run code (4) only when the
session actually requested dry-run; a stray STATUS_DRY_RUN_TRANSFER from a
hostile/buggy peer is now a logged protocol error (STATUS_ERROR) instead of
falling through to send file data and desync. Both normal send_single_file
callers handle rc == 4 explicitly as an abort.
A wire dry_run bit must not relax the destination-root precondition:
previously handlers skipped ensure_receive_root entirely in dry-run, so a
client could dry-run against a nonexistent/regular-file root a real session
rejects. Split the existence check (receive_root_exists, never creates)
from the create path and apply the precondition unconditionally: dry-run
runs the existence/directory check only, reports the failure, and creates
nothing (no --mkpath).
Also allow a `read only = yes` daemon module for a dry-run session (a
server-contacting dry-run IS a read-only wire operation) while still
refusing it for real writes, and update the stale read-only comments.
--dry-run now handshakes with a remote/daemon receiver and reports what
WOULD transfer/skip based on receiver state, mutating nothing on either
side.
- Serialize Config.dry_run into the wire config frame and append
STATUS_DRY_RUN_TRANSFER to the status enum (no renumbering); bump
PROTOCOL_VERSION/CMake VERSION/CHANGELOG/golden wire to 2.21.0.
- Receiver: receive_incremental_check_ex runs the normal read-only
decision and answers STATUS_OK (skip) or STATUS_DRY_RUN_TRANSFER
(would transfer) with no basis materialization/append/delta/full
transfer. All mutation sites are guarded by !dry_run: file store,
manifest deletes, --mkpath root creation, --delay-updates staging,
publication, directory-time application, and outcome acks.
- Client: send_dry_run_remote connects, sends the config, checks each
regular file and prints the would-transfer set + trailer; no file data
or delete manifest is sent. Plain local destinations keep the
client-side manifest.
Stamp host_last_use before publishing a bucket key and treat an unstamped
(last_use == 0) bucket as live, so a just-claimed bucket can no longer be
stolen by a concurrent reclaimer.
After a successful eviction CAS, re-scan for the interned key and, when an
earlier bucket already holds it, zero the duplicate's active count and
return the canonical bucket, preventing orphaned per-host counts and cap
overshoot under full-table concurrency.
Add a message-carrying EXPECT_FAIL primitive and use it for the daemon-conf
buffer-overflow guard, and add a fork-based test that records auth failures
from forked children and asserts the parent observes the shared lockout.
The config_receive_{basis,skip,idmap}_count helpers wrote the
peer-controlled int through the Config member before range-checking it.
An over-cap basis_count therefore left config->basis_count huge while
config->basis_dirs was still NULL; config_receive()'s error path then
called config_delete(), whose basis loop dereferenced NULL and crashed
the daemon before authentication.
Read each count into a local, validate, and only then assign, leaving the
member untouched on failure. config_delete() also guards the basis loop
with the array pointer as defense in depth.
Add a regression test that feeds over-cap basis/idmap/skip counts and
asserts rejection without crashing, plus a direct config_delete() check
on the partial (count set, array NULL) state.
Address low-severity review findings on the X-macro config refactor:
1. The golden test only hashed config_send_wire_block(), so a
receive-side KIND that reads a different width/order could still
round-trip symmetrically. Add test_config_wire_golden_receive():
capture the same hash-pinned 633-byte frame and feed it through
config_receive(), asserting every field (config_wire_equal) plus the
derived use_delta/use_xattrs bits and representative bounded kinds.
Add test_config_wire_receive_bounds() for bounds the symmetric
round-trip cannot reach: an out-of-range BOOL (hand-built frame),
RAW_MAXALLOC zero, a malformed STR_MODULE, an over-cap
INT_IDMAPCOUNT, and an out-of-range INT_IDENTITY chown_uid.
2. golden_config_populate() set long runs of booleans to all-1, so an
adjacent swap within a run produced identical bytes. Alternate the
boolean values and make the fixture receiver-valid (chmod grammar
"u=rwx,go=rx" is the same 11 bytes; delta_max_file_size inside the
bound). Re-pin the golden: len stays 633, hash is now
9160991280011164139 (computed, not guessed).
3. Document in config.h and client_cli.c that the CLI option tables
remain hand-maintained and are deliberately not generated from the
wire-field X-macro (client-only fields, flag/alias/negation
semantics). No CLI-table rewrite.
PROTOCOL_VERSION stays "2.20.0"; src/shared/config.c is untouched and
the wire bytes are unchanged apart from the fixture's own new values.
Every client on loopback shares the 127.0.0.1 identity, so counting them
against 'max connections per host' or the default-on auth lockout lets one
local client deny service to all the others (and makes a shared-NAT/proxy
address a natural DoS vector for remote clients). Use
utils_fd_peer_is_local (fail-closed) in the daemon gate to exempt a
provably local peer from the per-source cap and the auth lockout while
keeping the per-module and global caps. Remote peers are unchanged.
Document the shared-NAT/proxy identity limitation and the loopback
exemption in README/RSYNC_COMPAT/CHANGELOG, update the integration test to
assert the exemption, and fix the README 'auth failure delay' cap (5000,
not 60000).
The per-source host table only grew: once its fixed open-addressed table
filled, host_intern returned -1 and the per-host cap plus the shared auth
lockout silently failed open forever. Add a bounded-lifetime eviction
policy: track a per-bucket last-use time and, when no empty bucket exists,
atomically repurpose the first bucket that has no active connection and
either has an expired lockout or has been idle, resetting its counters.
Warn (rate-limited) on the genuine fail-open path.
A child SIGKILLed mid-registration could also leak a module/host count
because the parent only decremented on a REGISTERED slot. Make the slot
table the source of truth: after the SIGCHLD reap the parent recomputes
module_active[]/host_active[] from the surviving REGISTERED slots (atomics
only, async-signal-safe) so any leaked increment is erased.
Also clamp module_count to DAEMON_LIMITS_MAX_MODULES and use one helper
for the sizing/register host-tracking condition (a lockout threshold with
duration 0 is a no-op and must not intern hosts).
Add a NULL guard to protocol_release_memory_for_session so it no-ops like
the sibling session setters. Correct the Data.owner doc comment, which
implied a non-zero protocol_charge always has an owner; document that
owner may be NULL for uncharged/ownerless Data, that any such charge
falls back to the bound session, and that a charged Data must not outlive
its owning session. Note the lifetime contract on the release API too.
Extend tests/test_protocol.c to cover destroying a charged Data with no
session bound (the other half of the original bug) and to assert that
data_create/data_create_reserve start with owner == NULL and
protocol_charge == 0.
- pr-review: replace invalid 'tea pr comment' with 'tea comment' (the
former is not a tea subcommand)
- integrator: drop stray '-M' from client examples (-M is now
--remote-option and requires an argument), use the canonical pytest
integration command, and bump the CI image tag to v10
- test-writer: build fuzz targets via -DENABLE_FUZZ=ON instead of
hand-rolled -fsanitize flags; fix the fuzz binary path
- cmake-expert: document -DSANITIZER=undefined, which is now live in
CMakeLists.txt
- README: add --allow-unauthenticated to the plain-TCP server example,
use --preserve for metadata (not -M), and use the canonical
integration command
- AGENTS.md: use the canonical integration command
Document on utils_get_authorized_root_path() that the returned pointer is
borrowed and invalidated by the next authorized-root setter, that the fd
and path are not read atomically (non-reentrant), and that the fd remains
caller-owned. Add a matching single-threaded/set-before-threads note at
the accessor definitions in utils.c.
In server.c, drop the redundant utils_set_authorized_root(-1, NULL) after
a failed utils_set_authorized_root(): the setter already fail-closes the
state on allocation failure. The following close(root_fd) is unchanged.
The agent and skill definitions had drifted badly from the current
codebase and tooling, repeating the same class of bug as the benchmark
tool (references to nonexistent scripts and invented flags):
- Replace the removed `python3 test.py` with the real integration
command (`python3 -m pytest tests/integration/ -n 4 --dist=load
-m "not setpriv"`) across agents and skills.
- Fix `feature-scout`'s fabricated CLI flag list (--host, --server-mode,
--use-* etc.) using the authoritative src/client/usage.c flags.
- Fix `perf-analyst` benchmark flags (-m -c -> -j -z) and point at
benchmark/bench.py instead of stale numbers.
- Correct `code-explainer` (no getopt_long; --sendfile not -f) and
version drift in the release skill (1.1.0 -> 2.20.0).
- Replace GitHub/`gh` workflows with Gitea/`tea` (PRs target dev; issues
via tea; branch strategy updated in all agents).
- Use the built-in `-DSANITIZER=address|thread` CMake option instead of
hand-rolled -fsanitize flags.
- Add `-p 8080 --allow-unauthenticated` to plain-TCP server examples.
- Merge the redundant security-screener into security-auditor; drop the
duplicate (16 agents remain).
Repo hygiene: gitignore `root/` and `test_partial_install_tmp/`, remove
the empty leftover trees, delete the tracked scratch scripts tmux.sh and
to_one_file.py, and note the compile_commands.json symlink in README.
test_config_wire_golden() serializes a fully-populated Config through
config_send_wire_block() and pins the exact frame to len=633 and FNV-1a
hash 6163263374908258816, captured from the pre-X-macro implementation.
Any field reorder, resize or codec change fails the test.
test_config_wire_roundtrip_all_fields() serializes/deserializes a defaults
Config and a fully-populated Config over a socketpair and compares every
serialized field. The comparison is itself generated from
CONFIG_WIRE_FIELDS (one CONFIG_CMP_<KIND> per table entry), so a new table
entry automatically extends coverage; it cannot fall out of sync. It
normalizes the receiver's NULL/"" canonicalization, the max_alloc server
clamp and the derived use_delta/use_xattrs bits.
Every Config field that crosses the wire was declared in up to six places
(struct member, config_set_defaults, send_*, receive_*, and the two CLI
option tables) and could drift silently. Add CONFIG_WIRE_FIELDS in
config.h: one ordered per-segment table where each serialized field is
declared once with its C type, default and wire codec (KIND).
config.h now expands the table to declare the struct members;
config_set_defaults() expands it to assign the defaults; and
config_send_wire_block()/config_receive() expand the per-segment lists to
emit/consume the frame. The per-segment function names, call order and
segment boundaries are preserved exactly.
Fields with genuinely custom logic keep dedicated helpers but are still
declared once in the table: the protocol-version handshake (HEADER), daemon
SCRAM auth (STR_REDACTED_AUTH), the daemon module name (STR_MODULE), the
repeated count+array blocks (BLOCK_SKIP_SUFFIXES/BLOCK_BASIS/BLOCK_IDMAP),
--copy-as presence/ids (COPY_AS_*), and the derived --delta / use_xattrs
bits (DERIVED_DELTA, BOOL_XATTR_DERIVE). The version field remains a
special header (validated before any other field is parsed) and is sent by
config_send_wire_block() explicitly.
No public field is renamed and PROTOCOL_VERSION stays "2.20.0". Because
the struct declaration order is no longer the wire order, the wire order is
now enforced solely by the table and by a byte-exact golden test
(follow-up commit). Add config_send_wire_block() so that test can hash the
frame body without the STATUS_OK handshake.
Wire the shared registry into the accept loop (parent claims a slot before
fork, blocks SIGCHLD across fork+pid publication, and reclaims the dead
child's slot from the SIGCHLD handler so per-module/per-source counts are
released even on SIGKILL). The connection child records the selected module
and normalized peer IP once the config frame names them: an over-cap module
or source is refused at the config gate with an audit log, and a source
that exceeded the auth-failure threshold is refused before a SCRAM
challenge (the counter is shared across children and cleared on success).
The existing global cap and host ACLs are untouched.
Add global keys `max connections per host` (default 0 = unlimited),
`auth lockout threshold` (default 10, 0 disables) and
`auth lockout duration` (default 300 s, 0 disables). Module
`max connections` now accepts 0 as unlimited. Bound the number of
[module] sections (DAEMON_CONF_MAX_MODULES) so the shared registry's
per-module counter array stays fixed-size; absent keys keep their
defaults so old configs still load.
The daemon forks one child per accepted connection, so per-module and
per-source accounting must live in state shared across the children. Add a
fixed-size registry carved from an anonymous shared mapping
(mmap(MAP_SHARED|MAP_ANONYMOUS)) created before the accept loop: a slot
lifecycle (FREE/CLAIMED/REGISTERED) with parent claim/reclaim and a
lock-free, open-addressed per-source table for the per-host occupancy and
the shared auth-failure counter. C11 atomics only; no pthread locks across
fork.
Unit tests cover slot exhaustion, the module/host caps, pid reclaim and
fork-shared visibility.
The benchmark tool used stale rsync-style spellings that map to
different FastSync options, so it never enabled the features it
claimed to measure:
-c -> --checksum (not compression)
-m -> --prune-empty-dirs (not multithreading)
-s -> --secluded-args, a no-op (not chunk serialization)
-f -> --filter, needs an argument (not sendfile)
Replace them with the real flags (-z, -j, --chunk-serialization,
--sendfile), force CMAKE_BUILD_TYPE=Release, route informational
output to stderr so --output json emits valid JSON, surface
client/rsync failures instead of silently dropping them, and widen
the results table for the longer config names. Update the benchmark
skill to match (correct flags, server invocation, and replace the
nonexistent test.py --full with benchmark/bench.py).
Data charged against a ProtocolSession kept only the charge amount, so
data_destroy released it from whatever session was thread-locally bound
at destroy time. Destroying a received Data on another thread, after the
session was unbound, or while a different session was bound leaked the
originating session's budget and underflowed the other's.
Add Data.owner, set it whenever protocol_receive_data_limited charges a
session, and have data_destroy release against that owner directly via
the newly-exported protocol_release_memory_for_session. Uncharged Data
(owner NULL) keeps the previous bound-session fallback.
Add a unit test proving a Data acquired on session A is released to A
even when unrelated session B is bound at destroy time.
Break the ~700-line parse_args god function into cohesive static helpers
grouped by concern: output controls, pre-negation, range/time options, the
OPTION_TABLE dispatcher, flag/meta handlers, IO/network options, filter and
logging options, checksum/socket options, remote/basis/identity options,
positional handling, and a final lowering step.
A file-local CliParseCtx carries the config, cursor, positional buffers and
the mutable parse flags, so each handler stays focused. The dispatcher calls
the handlers in the original recognition order and preserves the exact
return contract (0/1/negative), error messages, log levels and control flow.
Behavior preserved; no functional changes.
Type Config.super_mode as SuperMode (a proper C enum) instead of a bare
int. The wire boundary still carries the mode as an int: send casts the
enum explicitly and receive reads a temporary int, validates the
AUTO..OFF range, then casts. Emitted bytes and accepted values are
unchanged. ModuleGateContext.super_mode_override keeps its -1 sentinel
as int with an explicit cast at the apply site.
Behavior preserved.
Deduplicate the repeated directory-time capture gate
(`config->use_metadata && !config->omit_dir_times`) used by the
sender-side (multiprocessing.c) and receiver-side (receiver.c) sinks
into a single predicate declared next to the DirTimeList machinery in
file_receive.h and defined in file_receive.c.
Behavior preserved: identical short-circuit condition and semantics,
no signature or protocol changes.
- Protocol version 2.19.0 (SCRAM-SHA-256 daemon auth replacing the replayable digest)
- Salted PBKDF2 verifier store + --hash-credentials; legacy store hard-rejected
- Persistent anti-enumeration dummy key (<store>.dummykey)
- Verified TLS / opted-in loopback transport required for auth modules
- Secret wiping; carried-over hardening from the security phases
- Add CHANGELOG.md and set the CMake project version
- server gate: the --allow-unauthenticated loopback allowance now requires
an actual plaintext connection (!gate_ctx->ssl), so a loopback TLS client
whose cert fails the --client-cn check is refused before any SCRAM
challenge instead of falling through the plaintext opt-in. Keep the
invalid-fd guard as belt-and-braces (unreachable after the policy check).
- test: rewrote test_wrong_client_cn_refused_before_auth_challenge to run
deterministically over 127.0.0.1 with --tls + --allow-unauthenticated and
a CA-valid wrong-CN client cert, asserting the gate refusal log and an
unchanged module tree (no skip).
- docs: --client-cn is mandatory with --tls; dummykey sidecar is secret
material; document all transient-fallback reasons; qualify
--allow-unauthenticated in README and --help so it cannot read as
permitting remote plaintext auth.
- credentials.h: drop stale restrictive-umask claim (fchmod forces exact
0600; only create/write/fsync/link/fchmod failure degrades to ephemeral).
- Make the atomic-publish temp name unpredictable by appending 16 random
hex chars to the pid, so a leftover/planted temp cannot be targeted.
- On EEXIST, unlink the stale temp and retry the O_EXCL create once
(bounded), so a crash leftover or reused pid cannot silently defeat
sidecar persistence.
- fchmod the temp fd to 0600 after creation (umask can clear owner bits)
and treat failure as a create failure, so the published sidecar is
always exactly 0600.
- Clarify comments: the sidecar requires exact 0600 while the store and
password files only reject group/other bits.
- Add a unit test that a restrictive umask still yields an exact 0600
sidecar; clean random-suffixed temps in tests.
utils_fd_peer_is_local now returns true only when getpeername SUCCEEDS and the
peer address classifies as loopback. A non-socket descriptor (pipe/socketpair)
or any getpeername error is NOT local, so the daemon auth gate fails closed
instead of treating an untestable --stdio pipe as trusted (daemon auth modules
are --daemon-only and the stdio path never loads a daemon config).
server_module_gate now requires --allow-unauthenticated for the loopback
plaintext auth path: a plaintext loopback connection without the operator
opt-in is refused at the config gate BEFORE server_auth_handshake, so no SCRAM
challenge is sent. Remote peers still require verified TLS regardless of the
flag; the handler keeps its defense-in-depth checks.
Docs state the exact policy (verified TLS with matching --client-cn, or
operator-opted-in loopback plaintext), drop the SSH/stdio auth-transport claim
(they are daemon-only), and add the loopback trust-boundary relay caveat and
the CN-only (no SAN) residual. Adds a unit-test negative for pipe/socketpair
and an integration test where a relay observes no challenge when the flag is
absent.
Address review findings on the persistent dummy-key sidecar:
- Publish atomically: write a private same-directory temp file
(<store>.dummykey.tmp.<pid>, 0600), fsync, then link(2) into place;
fsync the containing directory and drop the temp name. A concurrent
starter can no longer observe a zero/partial sidecar and fail closed.
On EEXIST adopt the winner's sidecar; otherwise warn and use a
transient ephemeral key.
- Harden the read path (initial and EEXIST-adopt) with
O_RDONLY|O_NOFOLLOW|O_NONBLOCK|O_CLOEXEC: reject planted symlinks
(ELOOP fails closed) and never block on a planted FIFO.
- Require the exact owner-only mode (st_mode & 07777) == 0600 and make
the rejection message truthful.
- Report a clear "short write" instead of a stale strerror(errno) when
write() returns 0.
- Document the artifact and its creation-failure caveat (FIFO store
path, read-only filesystem, missing directory) in README.md and
RSYNC_COMPAT.md.
- Tests: known-key sidecar adoption (dummy salt KAT + reload), symlink
rejection, and the exact-0600 rule (0400 now rejected).
Daemon modules that declare 'auth users' no longer accept credentials over a
remote plaintext connection: server_module_gate refuses at the config gate,
before any SCRAM challenge is sent, unless the connection is verified TLS with
a client certificate matching --client-cn, or a local/SSH transport (loopback
TCP peer or the --stdio pipe). --allow-unauthenticated does not relax this.
The TLS client-CN comparison now uses credentials_secure_equal (S2). Clients
sending --password-file to a non-loopback daemon must use --tls; validate_config
rejects the plaintext case before any network I/O.
Adds utils_sockaddr_is_loopback / utils_fd_peer_is_local / utils_host_is_loopback
helpers with unit tests, a client validation unit test, and integration tests
for the client-side plaintext rejection and the wrong-CN gate refusal.
The store-wide dummy key was regenerated on every credentials_load, so an
unknown user's dummy salt changed across daemon restarts while a real user's
stored salt stayed stable -- a restart-gated username-enumeration oracle.
Persist the 32-byte key in a 0600 <store>.dummykey sidecar next to the
credential store. An absent sidecar is created with O_EXCL and fsynced; a
present sidecar is read only when it is an owner-only regular file of exactly
32 bytes (otherwise the load fails closed). If the sidecar cannot be created
(read-only mount, missing directory) fall back to a transient per-run key with
a warning. A NULL store path keeps the key ephemeral.