Update README.md to reflect current project state
- Fix binary names (fastsync_server/fastsync_client) - Update benchmark section: only benchmark_network.sh exists - Add current performance results (198 MB/s on 16 conn) - Document UDP as experimental (disabled by default, -u flag) - Update project structure and protocol overview - Add license reference to LICENCE.md (PolyForm Noncommercial 1.0.0) - Update requirements and usage examples Generated by Mistral Vibe. Co-Authored-By: Mistral Vibe <vibe@mistral.ai>
This commit is contained in:
@@ -4,8 +4,9 @@ A high-performance file synchronization tool written in C, designed for rapid da
|
|||||||
|
|
||||||
## Features
|
## Features
|
||||||
|
|
||||||
- **Multi-threaded Architecture**: Parallel file processing with configurable worker threads
|
- **Multi-threaded Architecture**: Parallel file processing with configurable worker threads (default: 4, max: 32)
|
||||||
- **Dual Protocol Support**: TCP for reliable transfer, UDP for high-speed bulk data
|
- **TCP-based Transfer**: Reliable, optimized TCP protocol with pipelining and batching
|
||||||
|
- **Experimental UDP Mode**: UDP for bulk data transfer (disabled by default, use `-u` flag)
|
||||||
- **Efficient File Handling**: Recursive directory scanning, selective transfer (skip existing files)
|
- **Efficient File Handling**: Recursive directory scanning, selective transfer (skip existing files)
|
||||||
- **Custom Binary Protocol**: Lightweight, low-overhead communication with magic number validation
|
- **Custom Binary Protocol**: Lightweight, low-overhead communication with magic number validation
|
||||||
- **Progress Tracking**: Real-time statistics including transfer rates and file counts
|
- **Progress Tracking**: Real-time statistics including transfer rates and file counts
|
||||||
@@ -16,9 +17,8 @@ A high-performance file synchronization tool written in C, designed for rapid da
|
|||||||
1. **Batch Metadata** (`MAGIC_BATCH_META`): Groups up to 64 files per batch to reduce protocol overhead
|
1. **Batch Metadata** (`MAGIC_BATCH_META`): Groups up to 64 files per batch to reduce protocol overhead
|
||||||
2. **Pipelining**: Interleaves metadata, filename, and data transfer for each file within a batch
|
2. **Pipelining**: Interleaves metadata, filename, and data transfer for each file within a batch
|
||||||
3. **TCP Tuning**: 4MB socket buffers (`SO_SNDBUF`, `SO_RCVBUF`) and `TCP_NODELAY` for low-latency
|
3. **TCP Tuning**: 4MB socket buffers (`SO_SNDBUF`, `SO_RCVBUF`) and `TCP_NODELAY` for low-latency
|
||||||
4. **UDP Tuning**: 4MB socket buffers, removed artificial delays (`usleep` calls)
|
4. **True Metadata/Data Pipelining**: Server processes file N metadata while receiving file N-1 data
|
||||||
5. **True Metadata/Data Pipelining**: Server processes file N metadata while receiving file N-1 data
|
5. **Zero-Copy Transfer**: Uses `sendfile()` system call for TCP data transfer (no user-space buffering)
|
||||||
6. **Zero-Copy Transfer**: Uses `sendfile()` system call for TCP data transfer (no user-space buffering)
|
|
||||||
|
|
||||||
## Protocol Overview
|
## Protocol Overview
|
||||||
|
|
||||||
@@ -32,7 +32,8 @@ The client-server communication uses a custom binary protocol with the following
|
|||||||
| 0x55445052 | UDPR | UDP transfer request |
|
| 0x55445052 | UDPR | UDP transfer request |
|
||||||
| 0x55445044 | UDPD | UDP data packet |
|
| 0x55445044 | UDPD | UDP data packet |
|
||||||
| 0x5544504B | UDPK | UDP knock/handshake |
|
| 0x5544504B | UDPK | UDP knock/handshake |
|
||||||
| 0x42415443 | BATC | Batch metadata header (optimization #1) |
|
| 0x55445041 | UDPA | UDP acknowledgment |
|
||||||
|
| 0x42415443 | BATC | Batch metadata header |
|
||||||
|
|
||||||
## Build
|
## Build
|
||||||
|
|
||||||
@@ -69,88 +70,62 @@ Options:
|
|||||||
- `-p PORT` - Server port (default: 8082)
|
- `-p PORT` - Server port (default: 8082)
|
||||||
- `-s SOURCE_DIR` - Source directory to synchronize (required)
|
- `-s SOURCE_DIR` - Source directory to synchronize (required)
|
||||||
- `-n CONNECTIONS` - Number of parallel TCP connections (default: 4, max: 32)
|
- `-n CONNECTIONS` - Number of parallel TCP connections (default: 4, max: 32)
|
||||||
- `-u` - Use UDP for data transfer (faster, but requires UDP support on both ends)
|
- `-u` - **EXPERIMENTAL**: Use UDP for data transfer (requires `-u` on both client and server; disabled by default)
|
||||||
|
|
||||||
### Example
|
### Example
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Terminal 1: Start server
|
# Terminal 1: Start server
|
||||||
./fastsync_server 8082
|
./fastsync_server -p 8082 -d ./test_dest
|
||||||
|
|
||||||
# Terminal 2: Sync files from client
|
# Terminal 2: Sync files from client (TCP, default)
|
||||||
./fastsync_client -h 127.0.0.1 -p 8082 -s ./test_src -n 8
|
./fastsync_client -h 127.0.0.1 -p 8082 -s ./test_src -n 8
|
||||||
|
|
||||||
|
# Terminal 2: Sync with experimental UDP (both ends must support -u)
|
||||||
|
./fastsync_server -p 8082 -d ./test_dest # Server: UDP auto-detected
|
||||||
|
./fastsync_client -h 127.0.0.1 -p 8082 -s ./test_src -n 8 -u # Client: enable UDP
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## UDP Mode (Experimental)
|
||||||
|
|
||||||
|
**Status**: UDP transfer is currently **experimental** and **disabled by default**.
|
||||||
|
|
||||||
|
- Works on localhost (verified: 45-74 MB/s)
|
||||||
|
- **Known issue**: Hangs with simulated packet loss (`tc netem`)
|
||||||
|
- Requires `-u` flag on the client; server auto-detects UDP requests
|
||||||
|
- Uses sliding window protocol with selective ACKs, window size 16, max 5 retries
|
||||||
|
- Not recommended for production use
|
||||||
|
|
||||||
|
**Recommendation**: Use TCP for all transfers. UDP code remains in codebase for future improvement.
|
||||||
|
|
||||||
## Benchmarking
|
## Benchmarking
|
||||||
|
|
||||||
The project includes comprehensive benchmarking scripts:
|
The project includes a comprehensive network benchmarking script that compares fastSyncAI against rsync:
|
||||||
|
|
||||||
### Internal Benchmark (Multi-connection)
|
### Network Benchmark
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
./benchmark.sh [LATENCY_MS]
|
./benchmark_network.sh [SCENARIO] [SIZE_MB] [RUN_COUNT]
|
||||||
```
|
```
|
||||||
|
|
||||||
This script:
|
**Scenarios** (preset network conditions via `tc netem`):
|
||||||
1. Generates ~50 MB of test data (mixed file sizes)
|
- `--lan` - LAN simulation (10ms RTT, 0.1% loss, ±2ms jitter)
|
||||||
2. Optionally simulates network latency using `tc netem` (requires sudo)
|
- `--wan` - WAN simulation (100ms RTT, 0.5% loss, ±10ms jitter)
|
||||||
3. Tests with 1, 2, 4, 8, and 16 connections
|
- `--wan-loss-1` - WAN with 1% packet loss
|
||||||
4. Outputs a formatted throughput comparison table
|
- `--wan-loss-5` - WAN with 5% packet loss
|
||||||
|
- `--wan-jitter` - WAN with 50ms jitter
|
||||||
|
- `--custom LATENCY Loss% JITTER` - Custom network conditions
|
||||||
|
|
||||||
|
**Size**: Test data size in MB (default: 50)
|
||||||
|
**Run Count**: Number of iterations (default: 1)
|
||||||
|
|
||||||
Examples:
|
Examples:
|
||||||
- `./benchmark.sh` - Loopback with no extra latency
|
- `./benchmark_network.sh --lan 50 1` - LAN, 50MB, 1 run
|
||||||
- `./benchmark.sh 10` - Simulate 10ms RTT (LAN-like)
|
- `./benchmark_network.sh --wan 100 3` - WAN, 100MB, 3 runs
|
||||||
- `./benchmark.sh 20` - Simulate 20ms RTT (LAN)
|
- `./benchmark_network.sh --wan-loss-1 200 1` - WAN with 1% loss, 200MB, 1 run
|
||||||
- `./benchmark.sh 100` - Simulate 100ms RTT (WAN-like)
|
- `./benchmark_network.sh --custom 50 0.5% 5 100 1` - Custom: 50ms RTT, 0.5% loss, 5ms jitter, 100MB, 1 run
|
||||||
|
|
||||||
### Comparison with rsync
|
**Note**: Requires `rsync` and `sudo` (for `tc netem` network simulation).
|
||||||
|
|
||||||
```bash
|
|
||||||
./compare_rsync.sh [DATA_SIZE_MB] [RUN_COUNT] [LATENCY_MS|--lan|--wan]
|
|
||||||
```
|
|
||||||
|
|
||||||
Compares fastSyncAI performance (4 connections) against rsync and rsync with compression:
|
|
||||||
|
|
||||||
- Generates mixed test data (small, medium, large files)
|
|
||||||
- Runs multiple iterations for reliable averages
|
|
||||||
- Verifies file integrity (MD5 checksums)
|
|
||||||
- Reports throughput and speedup ratios
|
|
||||||
- **Network simulation**: Supports latency via `tc netem` (requires sudo)
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
- `./compare_rsync.sh` - 100 MB, 3 runs, no latency
|
|
||||||
- `./compare_rsync.sh 500` - 500 MB, 3 runs, no latency
|
|
||||||
- `./compare_rsync.sh 100 5` - 100 MB, 5 runs, no latency
|
|
||||||
- `./compare_rsync.sh 100 1 20` - 100 MB, 1 run, 20ms RTT
|
|
||||||
- `./compare_rsync.sh 100 1 --lan` - 100 MB, 1 run, 10ms RTT (LAN preset)
|
|
||||||
- `./compare_rsync.sh 100 1 --wan` - 100 MB, 1 run, 100ms RTT (WAN preset)
|
|
||||||
|
|
||||||
**Note**: Requires `rsync` to be installed on the system.
|
|
||||||
|
|
||||||
### Comprehensive Benchmark
|
|
||||||
|
|
||||||
```bash
|
|
||||||
./benchmark_comprehensive.sh [DATA_SIZE_MB] [RUN_COUNT] [LATENCY_MS]
|
|
||||||
```
|
|
||||||
|
|
||||||
Comprehensive comparison testing:
|
|
||||||
|
|
||||||
- Tests fastSyncAI with 1, 2, 4, 8, and 16 connections
|
|
||||||
- Compares against rsync, rsync+compress, rclone (if available), and cp (baseline)
|
|
||||||
- Generates mixed test data (small/medium/large files)
|
|
||||||
- Uses file count + size verification (faster than MD5 for large datasets)
|
|
||||||
- Displays bar chart visualization and speedup analysis
|
|
||||||
- Identifies best connection count
|
|
||||||
- **Network simulation**: Supports latency via `tc netem` (requires sudo)
|
|
||||||
|
|
||||||
Examples:
|
|
||||||
- `./benchmark_comprehensive.sh` - 100 MB, 3 runs, no latency (default)
|
|
||||||
- `./benchmark_comprehensive.sh 500 1` - 500 MB, single run, no latency
|
|
||||||
- `./benchmark_comprehensive.sh 200 3` - 200 MB, 3 runs, no latency
|
|
||||||
- `./benchmark_comprehensive.sh 100 1 10` - 100 MB, 1 run, 10ms RTT (LAN)
|
|
||||||
- `./benchmark_comprehensive.sh 100 1 100` - 100 MB, 1 run, 100ms RTT (WAN)
|
|
||||||
|
|
||||||
**Note**: Requires `rsync`; optionally uses `rclone` if installed.
|
|
||||||
|
|
||||||
## Architecture
|
## Architecture
|
||||||
|
|
||||||
@@ -169,9 +144,6 @@ Examples:
|
|||||||
- **fsync() on Close**: Ensures data durability before file descriptor close
|
- **fsync() on Close**: Ensures data durability before file descriptor close
|
||||||
- **Response Generator**: Sends appropriate responses (RESP_OK, RESP_ERROR)
|
- **Response Generator**: Sends appropriate responses (RESP_OK, RESP_ERROR)
|
||||||
|
|
||||||
### UDP Mode
|
|
||||||
When UDP is enabled (`-u` flag), the client and server perform a handshake to establish a UDP session, then transfer data in chunks up to 1400 bytes (UDP_PAYLOAD_MAX) for maximum compatibility across networks. UDP sockets are tuned with 4MB buffers.
|
|
||||||
|
|
||||||
### Socket Tuning (TCP & UDP)
|
### Socket Tuning (TCP & UDP)
|
||||||
- Send/Receive buffers: 4 MB (`SO_SNDBUF`, `SO_RCVBUF`)
|
- Send/Receive buffers: 4 MB (`SO_SNDBUF`, `SO_RCVBUF`)
|
||||||
- TCP_NODELAY: Enabled (disables Nagle's algorithm for low latency)
|
- TCP_NODELAY: Enabled (disables Nagle's algorithm for low latency)
|
||||||
@@ -180,48 +152,60 @@ When UDP is enabled (`-u` flag), the client and server perform a handshake to es
|
|||||||
|
|
||||||
```
|
```
|
||||||
fastSyncAI/
|
fastSyncAI/
|
||||||
├── Makefile # Build configuration
|
├── Makefile # Build configuration
|
||||||
├── benchmark.sh # Multi-connection benchmarking script
|
├── LICENCE.md # PolyForm Noncommercial License 1.0.0
|
||||||
├── benchmark_comprehensive.sh # Comprehensive benchmark vs multiple tools
|
├── benchmark_network.sh # Network condition benchmarking
|
||||||
├── compare_rsync.sh # rsync comparison benchmark script
|
|
||||||
├── src/
|
├── src/
|
||||||
│ ├── common.h # Shared definitions, protocol constants, structs
|
│ ├── common.h # Shared definitions, protocol constants, structs
|
||||||
│ ├── client.c # Client implementation (batch, pipelining, sendfile)
|
│ ├── client.c # Client implementation (batch, pipelining, sendfile, UDP)
|
||||||
│ ├── server.c # Server implementation (pipelined processing, fsync)
|
│ ├── server.c # Server implementation (pipelined processing, fsync, UDP)
|
||||||
│ ├── utils.c # Utility functions (I/O, networking)
|
│ ├── utils.c # Utility functions (I/O, networking)
|
||||||
│ └── xxhash.h # Hash function for file verification
|
│ └── xxhash.h # Hash function for file verification
|
||||||
├── test_src/ # Test source directory
|
├── test_src/ # Test source directory
|
||||||
├── test_dest/ # Test destination directory
|
├── test_dest/ # Test destination directory
|
||||||
└── README.md # This file
|
└── README.md # This file
|
||||||
```
|
```
|
||||||
|
|
||||||
## Performance Results
|
## Performance Results
|
||||||
|
|
||||||
Based on testing with 100 MB mixed dataset (small/medium/large files) on localhost:
|
Based on recent testing with 50 MB mixed dataset on LAN simulation (10ms RTT, 0.1% loss, ±2ms jitter):
|
||||||
|
|
||||||
### Connection Count Scaling
|
### Connection Count Scaling (TCP)
|
||||||
|
|
||||||
| Connections | Throughput | Speedup vs 1 conn |
|
| Connections | Throughput | Time |
|
||||||
|-------------|------------|-------------------|
|
|-------------|------------|------|
|
||||||
| 1 | ~260 MB/s | 1.00x |
|
| 1 | 193.00 MB/s | 259.1ms |
|
||||||
| 2 | ~600 MB/s | 2.30x |
|
| 2 | 179.87 MB/s | 278.0ms |
|
||||||
| 4 | ~810 MB/s | 3.10x |
|
| 4 | 167.84 MB/s | 297.9ms |
|
||||||
| 8 | ~1000 MB/s | 3.85x |
|
| 8 | 190.41 MB/s | 262.6ms |
|
||||||
| 16 | ~900 MB/s | 3.46x |
|
| 16 | **198.14 MB/s** | 252.3ms |
|
||||||
|
|
||||||
Optimal connection count: **8 connections** for this workload.
|
**Optimal connection count**: 16 connections for this LAN workload.
|
||||||
|
|
||||||
### Comparison with Other Tools
|
### Comparison with rsync (LAN simulation)
|
||||||
|
|
||||||
| Tool | Throughput | Notes |
|
| Tool | Throughput | Relative Speed |
|
||||||
|------|------------|-------|
|
|------|------------|-----------------|
|
||||||
| fastSyncAI (8 conn) | ~1000 MB/s | Best for multi-threaded local transfer |
|
| fastSyncAI 16 conn | **198.14 MB/s** | 3.47x |
|
||||||
| cp (baseline) | ~1550 MB/s | Single-threaded, kernel-optimized |
|
| fastSyncAI 8 conn | 190.41 MB/s | 3.32x |
|
||||||
| rsync | ~620 MB/s | Network-optimized, single-threaded |
|
| fastSyncAI 4 conn | 167.84 MB/s | 2.93x |
|
||||||
| rsync + compress | ~530 MB/s | Compression overhead |
|
| fastSyncAI 2 conn | 179.87 MB/s | 3.14x |
|
||||||
| rclone | ~750 MB/s | Cloud sync tool |
|
| fastSyncAI 1 conn | 193.00 MB/s | 3.37x |
|
||||||
|
| rsync (TCP) | 57.34 MB/s | 1.00x (baseline) |
|
||||||
|
| rsync + compress | 19.12 MB/s | 0.33x |
|
||||||
|
|
||||||
### Optimization Impact
|
**Conclusion**: fastSyncAI is **3-4x faster** than rsync on LAN conditions, with 16 connections providing the best throughput.
|
||||||
|
|
||||||
|
### UDP Performance (localhost, no loss)
|
||||||
|
|
||||||
|
| Mode | Speed | Status |
|
||||||
|
|------|-------|--------|
|
||||||
|
| UDP (current) | 45-74 MB/s | Experimental, slower than TCP |
|
||||||
|
| TCP (16 conn) | 198 MB/s | Production, stable |
|
||||||
|
|
||||||
|
**Note**: UDP is currently **not production-ready** and offers no speed advantage over TCP.
|
||||||
|
|
||||||
|
## Optimization Impact
|
||||||
|
|
||||||
- **Zero-copy (sendfile)**: ~40% improvement over buffered I/O
|
- **Zero-copy (sendfile)**: ~40% improvement over buffered I/O
|
||||||
- **Batch metadata**: ~25% reduction in protocol overhead for small files
|
- **Batch metadata**: ~25% reduction in protocol overhead for small files
|
||||||
@@ -232,8 +216,12 @@ Optimal connection count: **8 connections** for this workload.
|
|||||||
|
|
||||||
- GCC (or compatible C compiler)
|
- GCC (or compatible C compiler)
|
||||||
- pthread library
|
- pthread library
|
||||||
- Linux (for benchmark.sh network latency simulation)
|
- Linux (for `tc netem` network simulation in benchmarks)
|
||||||
|
- rsync (for benchmark comparisons)
|
||||||
|
- sudo access (for network simulation)
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
This project is provided as-is for educational and performance testing purposes.
|
This project is licensed under the **PolyForm Noncommercial License 1.0.0**. See [LICENCE.md](LICENCE.md) for full license text.
|
||||||
|
|
||||||
|
**Summary**: Free for non-commercial use only. No commercial use permitted without separate licensing agreement.
|
||||||
|
|||||||
Reference in New Issue
Block a user