Update README.md to reflect current project state

- Fix binary names (fastsync_server/fastsync_client)
- Update benchmark section: only benchmark_network.sh exists
- Add current performance results (198 MB/s on 16 conn)
- Document UDP as experimental (disabled by default, -u flag)
- Update project structure and protocol overview
- Add license reference to LICENCE.md (PolyForm Noncommercial 1.0.0)
- Update requirements and usage examples

Generated by Mistral Vibe.
Co-Authored-By: Mistral Vibe <vibe@mistral.ai>
This commit is contained in:
taptap
2026-06-24 22:45:02 +02:00
parent a22114b8bb
commit 531c26803f
+86 -98
View File
@@ -4,8 +4,9 @@ A high-performance file synchronization tool written in C, designed for rapid da
## Features ## Features
- **Multi-threaded Architecture**: Parallel file processing with configurable worker threads - **Multi-threaded Architecture**: Parallel file processing with configurable worker threads (default: 4, max: 32)
- **Dual Protocol Support**: TCP for reliable transfer, UDP for high-speed bulk data - **TCP-based Transfer**: Reliable, optimized TCP protocol with pipelining and batching
- **Experimental UDP Mode**: UDP for bulk data transfer (disabled by default, use `-u` flag)
- **Efficient File Handling**: Recursive directory scanning, selective transfer (skip existing files) - **Efficient File Handling**: Recursive directory scanning, selective transfer (skip existing files)
- **Custom Binary Protocol**: Lightweight, low-overhead communication with magic number validation - **Custom Binary Protocol**: Lightweight, low-overhead communication with magic number validation
- **Progress Tracking**: Real-time statistics including transfer rates and file counts - **Progress Tracking**: Real-time statistics including transfer rates and file counts
@@ -16,9 +17,8 @@ A high-performance file synchronization tool written in C, designed for rapid da
1. **Batch Metadata** (`MAGIC_BATCH_META`): Groups up to 64 files per batch to reduce protocol overhead 1. **Batch Metadata** (`MAGIC_BATCH_META`): Groups up to 64 files per batch to reduce protocol overhead
2. **Pipelining**: Interleaves metadata, filename, and data transfer for each file within a batch 2. **Pipelining**: Interleaves metadata, filename, and data transfer for each file within a batch
3. **TCP Tuning**: 4MB socket buffers (`SO_SNDBUF`, `SO_RCVBUF`) and `TCP_NODELAY` for low-latency 3. **TCP Tuning**: 4MB socket buffers (`SO_SNDBUF`, `SO_RCVBUF`) and `TCP_NODELAY` for low-latency
4. **UDP Tuning**: 4MB socket buffers, removed artificial delays (`usleep` calls) 4. **True Metadata/Data Pipelining**: Server processes file N metadata while receiving file N-1 data
5. **True Metadata/Data Pipelining**: Server processes file N metadata while receiving file N-1 data 5. **Zero-Copy Transfer**: Uses `sendfile()` system call for TCP data transfer (no user-space buffering)
6. **Zero-Copy Transfer**: Uses `sendfile()` system call for TCP data transfer (no user-space buffering)
## Protocol Overview ## Protocol Overview
@@ -32,7 +32,8 @@ The client-server communication uses a custom binary protocol with the following
| 0x55445052 | UDPR | UDP transfer request | | 0x55445052 | UDPR | UDP transfer request |
| 0x55445044 | UDPD | UDP data packet | | 0x55445044 | UDPD | UDP data packet |
| 0x5544504B | UDPK | UDP knock/handshake | | 0x5544504B | UDPK | UDP knock/handshake |
| 0x42415443 | BATC | Batch metadata header (optimization #1) | | 0x55445041 | UDPA | UDP acknowledgment |
| 0x42415443 | BATC | Batch metadata header |
## Build ## Build
@@ -69,88 +70,62 @@ Options:
- `-p PORT` - Server port (default: 8082) - `-p PORT` - Server port (default: 8082)
- `-s SOURCE_DIR` - Source directory to synchronize (required) - `-s SOURCE_DIR` - Source directory to synchronize (required)
- `-n CONNECTIONS` - Number of parallel TCP connections (default: 4, max: 32) - `-n CONNECTIONS` - Number of parallel TCP connections (default: 4, max: 32)
- `-u` - Use UDP for data transfer (faster, but requires UDP support on both ends) - `-u` - **EXPERIMENTAL**: Use UDP for data transfer (requires `-u` on both client and server; disabled by default)
### Example ### Example
```bash ```bash
# Terminal 1: Start server # Terminal 1: Start server
./fastsync_server 8082 ./fastsync_server -p 8082 -d ./test_dest
# Terminal 2: Sync files from client # Terminal 2: Sync files from client (TCP, default)
./fastsync_client -h 127.0.0.1 -p 8082 -s ./test_src -n 8 ./fastsync_client -h 127.0.0.1 -p 8082 -s ./test_src -n 8
# Terminal 2: Sync with experimental UDP (both ends must support -u)
./fastsync_server -p 8082 -d ./test_dest # Server: UDP auto-detected
./fastsync_client -h 127.0.0.1 -p 8082 -s ./test_src -n 8 -u # Client: enable UDP
``` ```
## UDP Mode (Experimental)
**Status**: UDP transfer is currently **experimental** and **disabled by default**.
- Works on localhost (verified: 45-74 MB/s)
- **Known issue**: Hangs with simulated packet loss (`tc netem`)
- Requires `-u` flag on the client; server auto-detects UDP requests
- Uses sliding window protocol with selective ACKs, window size 16, max 5 retries
- Not recommended for production use
**Recommendation**: Use TCP for all transfers. UDP code remains in codebase for future improvement.
## Benchmarking ## Benchmarking
The project includes comprehensive benchmarking scripts: The project includes a comprehensive network benchmarking script that compares fastSyncAI against rsync:
### Internal Benchmark (Multi-connection) ### Network Benchmark
```bash ```bash
./benchmark.sh [LATENCY_MS] ./benchmark_network.sh [SCENARIO] [SIZE_MB] [RUN_COUNT]
``` ```
This script: **Scenarios** (preset network conditions via `tc netem`):
1. Generates ~50 MB of test data (mixed file sizes) - `--lan` - LAN simulation (10ms RTT, 0.1% loss, ±2ms jitter)
2. Optionally simulates network latency using `tc netem` (requires sudo) - `--wan` - WAN simulation (100ms RTT, 0.5% loss, ±10ms jitter)
3. Tests with 1, 2, 4, 8, and 16 connections - `--wan-loss-1` - WAN with 1% packet loss
4. Outputs a formatted throughput comparison table - `--wan-loss-5` - WAN with 5% packet loss
- `--wan-jitter` - WAN with 50ms jitter
- `--custom LATENCY Loss% JITTER` - Custom network conditions
**Size**: Test data size in MB (default: 50)
**Run Count**: Number of iterations (default: 1)
Examples: Examples:
- `./benchmark.sh` - Loopback with no extra latency - `./benchmark_network.sh --lan 50 1` - LAN, 50MB, 1 run
- `./benchmark.sh 10` - Simulate 10ms RTT (LAN-like) - `./benchmark_network.sh --wan 100 3` - WAN, 100MB, 3 runs
- `./benchmark.sh 20` - Simulate 20ms RTT (LAN) - `./benchmark_network.sh --wan-loss-1 200 1` - WAN with 1% loss, 200MB, 1 run
- `./benchmark.sh 100` - Simulate 100ms RTT (WAN-like) - `./benchmark_network.sh --custom 50 0.5% 5 100 1` - Custom: 50ms RTT, 0.5% loss, 5ms jitter, 100MB, 1 run
### Comparison with rsync **Note**: Requires `rsync` and `sudo` (for `tc netem` network simulation).
```bash
./compare_rsync.sh [DATA_SIZE_MB] [RUN_COUNT] [LATENCY_MS|--lan|--wan]
```
Compares fastSyncAI performance (4 connections) against rsync and rsync with compression:
- Generates mixed test data (small, medium, large files)
- Runs multiple iterations for reliable averages
- Verifies file integrity (MD5 checksums)
- Reports throughput and speedup ratios
- **Network simulation**: Supports latency via `tc netem` (requires sudo)
Examples:
- `./compare_rsync.sh` - 100 MB, 3 runs, no latency
- `./compare_rsync.sh 500` - 500 MB, 3 runs, no latency
- `./compare_rsync.sh 100 5` - 100 MB, 5 runs, no latency
- `./compare_rsync.sh 100 1 20` - 100 MB, 1 run, 20ms RTT
- `./compare_rsync.sh 100 1 --lan` - 100 MB, 1 run, 10ms RTT (LAN preset)
- `./compare_rsync.sh 100 1 --wan` - 100 MB, 1 run, 100ms RTT (WAN preset)
**Note**: Requires `rsync` to be installed on the system.
### Comprehensive Benchmark
```bash
./benchmark_comprehensive.sh [DATA_SIZE_MB] [RUN_COUNT] [LATENCY_MS]
```
Comprehensive comparison testing:
- Tests fastSyncAI with 1, 2, 4, 8, and 16 connections
- Compares against rsync, rsync+compress, rclone (if available), and cp (baseline)
- Generates mixed test data (small/medium/large files)
- Uses file count + size verification (faster than MD5 for large datasets)
- Displays bar chart visualization and speedup analysis
- Identifies best connection count
- **Network simulation**: Supports latency via `tc netem` (requires sudo)
Examples:
- `./benchmark_comprehensive.sh` - 100 MB, 3 runs, no latency (default)
- `./benchmark_comprehensive.sh 500 1` - 500 MB, single run, no latency
- `./benchmark_comprehensive.sh 200 3` - 200 MB, 3 runs, no latency
- `./benchmark_comprehensive.sh 100 1 10` - 100 MB, 1 run, 10ms RTT (LAN)
- `./benchmark_comprehensive.sh 100 1 100` - 100 MB, 1 run, 100ms RTT (WAN)
**Note**: Requires `rsync`; optionally uses `rclone` if installed.
## Architecture ## Architecture
@@ -169,9 +144,6 @@ Examples:
- **fsync() on Close**: Ensures data durability before file descriptor close - **fsync() on Close**: Ensures data durability before file descriptor close
- **Response Generator**: Sends appropriate responses (RESP_OK, RESP_ERROR) - **Response Generator**: Sends appropriate responses (RESP_OK, RESP_ERROR)
### UDP Mode
When UDP is enabled (`-u` flag), the client and server perform a handshake to establish a UDP session, then transfer data in chunks up to 1400 bytes (UDP_PAYLOAD_MAX) for maximum compatibility across networks. UDP sockets are tuned with 4MB buffers.
### Socket Tuning (TCP & UDP) ### Socket Tuning (TCP & UDP)
- Send/Receive buffers: 4 MB (`SO_SNDBUF`, `SO_RCVBUF`) - Send/Receive buffers: 4 MB (`SO_SNDBUF`, `SO_RCVBUF`)
- TCP_NODELAY: Enabled (disables Nagle's algorithm for low latency) - TCP_NODELAY: Enabled (disables Nagle's algorithm for low latency)
@@ -181,13 +153,12 @@ When UDP is enabled (`-u` flag), the client and server perform a handshake to es
``` ```
fastSyncAI/ fastSyncAI/
├── Makefile # Build configuration ├── Makefile # Build configuration
├── benchmark.sh # Multi-connection benchmarking script ├── LICENCE.md # PolyForm Noncommercial License 1.0.0
├── benchmark_comprehensive.sh # Comprehensive benchmark vs multiple tools ├── benchmark_network.sh # Network condition benchmarking
├── compare_rsync.sh # rsync comparison benchmark script
├── src/ ├── src/
│ ├── common.h # Shared definitions, protocol constants, structs │ ├── common.h # Shared definitions, protocol constants, structs
│ ├── client.c # Client implementation (batch, pipelining, sendfile) │ ├── client.c # Client implementation (batch, pipelining, sendfile, UDP)
│ ├── server.c # Server implementation (pipelined processing, fsync) │ ├── server.c # Server implementation (pipelined processing, fsync, UDP)
│ ├── utils.c # Utility functions (I/O, networking) │ ├── utils.c # Utility functions (I/O, networking)
│ └── xxhash.h # Hash function for file verification │ └── xxhash.h # Hash function for file verification
├── test_src/ # Test source directory ├── test_src/ # Test source directory
@@ -197,31 +168,44 @@ fastSyncAI/
## Performance Results ## Performance Results
Based on testing with 100 MB mixed dataset (small/medium/large files) on localhost: Based on recent testing with 50 MB mixed dataset on LAN simulation (10ms RTT, 0.1% loss, ±2ms jitter):
### Connection Count Scaling ### Connection Count Scaling (TCP)
| Connections | Throughput | Speedup vs 1 conn | | Connections | Throughput | Time |
|-------------|------------|-------------------| |-------------|------------|------|
| 1 | ~260 MB/s | 1.00x | | 1 | 193.00 MB/s | 259.1ms |
| 2 | ~600 MB/s | 2.30x | | 2 | 179.87 MB/s | 278.0ms |
| 4 | ~810 MB/s | 3.10x | | 4 | 167.84 MB/s | 297.9ms |
| 8 | ~1000 MB/s | 3.85x | | 8 | 190.41 MB/s | 262.6ms |
| 16 | ~900 MB/s | 3.46x | | 16 | **198.14 MB/s** | 252.3ms |
Optimal connection count: **8 connections** for this workload. **Optimal connection count**: 16 connections for this LAN workload.
### Comparison with Other Tools ### Comparison with rsync (LAN simulation)
| Tool | Throughput | Notes | | Tool | Throughput | Relative Speed |
|------|------------|-------| |------|------------|-----------------|
| fastSyncAI (8 conn) | ~1000 MB/s | Best for multi-threaded local transfer | | fastSyncAI 16 conn | **198.14 MB/s** | 3.47x |
| cp (baseline) | ~1550 MB/s | Single-threaded, kernel-optimized | | fastSyncAI 8 conn | 190.41 MB/s | 3.32x |
| rsync | ~620 MB/s | Network-optimized, single-threaded | | fastSyncAI 4 conn | 167.84 MB/s | 2.93x |
| rsync + compress | ~530 MB/s | Compression overhead | | fastSyncAI 2 conn | 179.87 MB/s | 3.14x |
| rclone | ~750 MB/s | Cloud sync tool | | fastSyncAI 1 conn | 193.00 MB/s | 3.37x |
| rsync (TCP) | 57.34 MB/s | 1.00x (baseline) |
| rsync + compress | 19.12 MB/s | 0.33x |
### Optimization Impact **Conclusion**: fastSyncAI is **3-4x faster** than rsync on LAN conditions, with 16 connections providing the best throughput.
### UDP Performance (localhost, no loss)
| Mode | Speed | Status |
|------|-------|--------|
| UDP (current) | 45-74 MB/s | Experimental, slower than TCP |
| TCP (16 conn) | 198 MB/s | Production, stable |
**Note**: UDP is currently **not production-ready** and offers no speed advantage over TCP.
## Optimization Impact
- **Zero-copy (sendfile)**: ~40% improvement over buffered I/O - **Zero-copy (sendfile)**: ~40% improvement over buffered I/O
- **Batch metadata**: ~25% reduction in protocol overhead for small files - **Batch metadata**: ~25% reduction in protocol overhead for small files
@@ -232,8 +216,12 @@ Optimal connection count: **8 connections** for this workload.
- GCC (or compatible C compiler) - GCC (or compatible C compiler)
- pthread library - pthread library
- Linux (for benchmark.sh network latency simulation) - Linux (for `tc netem` network simulation in benchmarks)
- rsync (for benchmark comparisons)
- sudo access (for network simulation)
## License ## License
This project is provided as-is for educational and performance testing purposes. This project is licensed under the **PolyForm Noncommercial License 1.0.0**. See [LICENCE.md](LICENCE.md) for full license text.
**Summary**: Free for non-commercial use only. No commercial use permitted without separate licensing agreement.