<- MyBlog
30th September, 2026
8 min read
Completion based IO vs Readiness based IO
Reading files / Listening to sockets in userspace programs written for linux is done through system calls.
epoll is a readiness based event notification we have been using, io-uring came in linux
5.1, it works using shared memory mapped ring buffers between the userspace
and kernel-space. To do I/O,
"submit" an I/O request to the kernel, after some time, you
should expect a completion queue entry in your ring buffer.
epoll is stable and great for its use case. It will continued to be used for normal scenarios
io-uring allows advance workloads and fulfils custom needs.
We dispatch completion check in the turn function inside tokio's scheduler for io-uring operations.
to submit our submissions to the kernel we call io_uring_enter(2)
this lets us wait until some CQEs are ready to be seen.
For custom workloads where you have a thread in userspace that's whole job is to submit
submission queues can be enabled using Submission queue polling but it should be disabled by default
and has trade offs with higher cpu and memory usage.
on a Limactl VM since I'm on macos
[lima@lima-default tokio-io-uring]$ uname -a
Linux lima-default 6.19.10-300.fc44.aarch64 #1 SMP PREEMPT_DYNAMIC Wed Mar 25 17:45:07 UTC 2026 aarch64 GNU/Linux
Running strace on a rust tokio file reading project
#[tokio::main]
async fn main() -> std::io::Result<()> {
let path = "example.txt";
let mut file = File::create(path).await?;
file.write_all(b"Hello from Tokio!\n").await?;
file.flush().await?;
drop(file);
let mut file = File::open(path).await?;
let mut contents = String::new();
file.read_to_string(&mut contents).await?;
print!("{contents}");
Ok(())
}
Compile with RUSTFLAGS="--cfg tokio_unstable" cargo b --release and strace shows
[lima@lima-default tokio-io-uring]$ strace ./target/release/tokio-io-uring
// initalize io-uring in tokio
io_uring_setup(256, {flags=0, ... })
mmap(NULL, 16384, PROT_R, ... )
mmap(NULL, 16384, PROT_R, ...)
io_uring_register(6, IORIN ...)
// register fd to watch
epoll_ctl(5, EPOLL_CTL_ADD, 6, ... )
// enter kernel
io_uring_enter(6, 1, 0, 0, NULL, 128) = 1
futex(0xaaab0a14d3a0, ...) = 0
mmap(NULL, 2162688, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS|MAP_STACK, -1, 0) = 0xffffbcb83000
madvise(0xffffbcb83000, 65536, MADV_GUARD_INSTALL) = 0
rt_sigprocmask(SIG_BLOCK, ~[], [], 8) = 0
clone3({flags=CLONE_VM|CLONE_FS|CLONE_FILES|CLONE_SIGHAND|CLONE_THREAD|CLONE_SYSVSEM|CLONE_SETTLS|CLONE_PARENT_SETTID|CLONE_CHILD_CLEARTID, child_tid=0xffffbcd924e8, parent_tid=0xffffbcd92190, exit_signal=0, stack=0xffffbcb83000, stack_size=0x20e9a0, tls=0xffffbcd927e0} => {parent_tid=[18539]}, 88) = 18539
rt_sigprocmask(SIG_SETMASK, [], NULL, 8) = 0
futex(0xaaab0a144f90, FUTEX_WAKE_PRIVATE, 1) = 1
futex(0xaaab0a14d3a0, FUTEX_WAIT_BITSET_PRIVATE, 1, NULL, FUTEX_BITSET_MATCH_ANY) = 0
close(7) = 0
io_uring_enter(6, 1, 0, 0, NULL, 128) = 1
futex(0xaaab0a14a918, FUTEX_WAKE_PRIVATE, 1) = 1
futex(0xaaab0a14a298, FUTEX_WAKE_PRIVATE, 1) = 1
futex(0xaaab0a14d3a0, FUTEX_WAIT_BITSET_PRIVATE, 2, NULL, FUTEX_BITSET_MATCH_ANY) = 0
write(4, "\1\0\0\0\0\0\0\0", 8) = 8
write(1, "Hello from Tokio!\n", 18Hello from Tokio!) = 18
close(7) = 0
write(4, "\1\0\0\0\0\0\0\0", 8) = 8
In tokio we have managed to mix epoll which is a readiness based API and io-uring
which is a completion based API. Why so?
Tokio works on a layer of abstraction higher than mio. MIO is lightweight
a deterministic set of code that interfaces with different operating systems
and event notification mechanim supported by the OS. The tokio runtime Handle then keeps
an instance of Poll from MIO in the handle struct.
Completion based
APIs can still be built on top of poll/future abstractions but
io-uring doesn't need it. io-uring still supports polling file descriptors in poll style
using IORING_OP_POLL_ADD
io-uring addition today allows you to do non blocking async File I/O.
The tokio ecosystem is still deciding what the final io-uring API will look like. Windows had io-uring
equivalent support available. The windows specific code is mostly in mio
The API should accomdate features like provided buffers and registered buffers. They are key in supporting
zero copy send / recv methods on the sockets and O_DIRECT reads.
The tokio-rs/tokio-uring repository had some experiments regarding how that API should look like.
Provided buffers creating pre allocated vectors for the kernel to choose from to fill.
It looked roughly like this
old tokio-uring experiment
Most of the io-uring implementation has been moved to tokio.
There's also registered buffers which is different from provided buffers. You give the ownership of your file descriptor and your buffer until you reap the completion queue from the ring This creates a different API, this is how tokio's io-uring internals look like and interestingly, Compio's API
use compio::{fs::File, io::AsyncReadAtExt};
let file = File::open("Cargo.toml").await.unwrap();
let (read, buffer) = file
.read_to_end_at(Vec::with_capacity(1024), 0)
.await
.unwrap();
assert_eq!(read, buffer.len());
let buffer = String::from_utf8(buffer).unwrap();
println!("{}", buffer);
So let's get distracted a bit look at strace for the same equivalent program in Compio and tokio, migrating our code to compio looks like this
use compio::fs::File;
use compio::io::{AsyncReadExt, AsyncWrite, AsyncWriteExt};
#[compio::main]
async fn main() -> std::io::Result<()> {
let path = "example.txt";
let mut file = File::create(path).await?;
file.write_all(b"Hello from Tokio!\n").await?;
file.flush().await?;
drop(file);
let mut file = File::open(path).await?;
let mut contents = String::new();
file.read_to_string(&mut contents).await?;
print!("{contents}");
Ok(())
}
But it doesn't map so cleanly yet. There's an error
error[E0599]: the method `write_all` exists for struct `compio::compio_fs::File`, but its trait bounds were not satisfied
--> src/main.rs:9:10
|
9 | file.write_all(b"Hello from Tokio!\n").await?;
| ^^^^^^^^^
|
::: /Users/Work/.cargo/registry/src/index.crates.io-1949cf8c6b5b557f/compio-fs-0.12.1/src/file.rs:51:1
|
51 | pub struct File {
| --------------- doesn't satisfy `compio::compio_fs::File: AsyncWriteExt` or `compio::compio_fs::File: AsyncWrite`
|
= note: the following trait bounds were not satisfied:
`compio::compio_fs::File: AsyncWrite`
which is required by `compio::compio_fs::File: AsyncWriteExt`
Compio's file only implement methods for fixed read and writes, I will need a std cursor abstraction to read it with io-uring. Tokio does a probe read similar to std::fs does.
We need to give the ownership of the buffer and file desciptor to the runtime and the runtime will return us back the owned values once corresponding completion queue entry is reaped
use std::io::Cursor;
use compio::buf::BufResult;
use compio::fs::File;
use compio::io::{AsyncReadExt, AsyncWrite, AsyncWriteExt};
#[compio::main]
async fn main() -> std::io::Result<()> {
let path = "example.txt";
let file = File::create(path).await?;
let mut file = Cursor::new(file);
let BufResult(result, _) =
file.write_all(b"Hello from Compio!\n").await;
result?;
file.flush().await?;
drop(file);
let file = File::open(path).await?;
let mut file = Cursor::new(file);
let BufResult(result, contents) =
file.read_to_string(String::new()).await;
result?;
print!("{contents}");
Ok(())
}
What's the strace looking like ?
io_uring_setup(1024, .. ) = 4
mmap(NULL, 65536, PROT_READ|PROT_WRITE, MAP_SHARED|MAP_POPULATE, 4, 0x10000000) = 0xffffb1a60000
mmap(NULL, 36928, PROT_READ|PROT_WRITE, MAP_SHARED|MAP_POPULATE, 4, 0) = 0xffffb1c75000
madvise(0xffffb1c75000, 36928, MADV_DONTFORK) = 0
madvise(0xffffb1a60000, 65536, MADV_DONTFORK) = 0
io_uring_setup(2, ....) = 0
munmap(0xffffb1cb3000, 136) = 0
munmap(0xffffb1cb4000, 128) = 0
close(5) = 0
io_uring_enter(4, 2, 1, IORING_ENTER_GETEVENTS|IORING_ENTER_NO_IOWAIT, NULL, 128) = 2
io_uring_enter(4, 1, 1, IORING_ENTER_GETEVENTS|IORING_ENTER_NO_IOWAIT, NULL, 128) = 1
close(5) = 0
io_uring_enter(4, 1, 1, IORING_ENTER_GETEVENTS|IORING_ENTER_NO_IOWAIT, NULL, 128) = 1
io_uring_enter(4, 1, 1, IORING_ENTER_GETEVENTS|IORING_ENTER_NO_IOWAIT, NULL, 128) = 1
io_uring_enter(4, 1, 1, IORING_ENTER_GETEVENTS|IORING_ENTER_NO_IOWAIT, NULL, 128) = 1
write(1, "Hello from Compio!\n", 19Hello from Compio!
) = 19
close(5) = 0
It has more io_uring_enter calls but less system calls overall.
Tokio's file read implements tricks similar to rust's std file read method to probe read a file fully
and allocate more space for it if needed.
This is because we need to provide the number of bytes we want to read for
each io-uring submission and wait for that read request for let's say N bytes is completed, release
that memory and give it back to the caller.
So how can we make tokio's mulithreaded work stealing runtime work with io-uring cleanly? A big problem in tokio today lies in syncronising the bookeeping of submission queues and completion queue in the poll implementation of each io-uring operation supported by tokio
Benchmarking your runtime implementation is not as simple as running a linux VM and running criteron benchmarks. Should benchmark on real databases / servers workloads. This makes it hard to do regression testing for runtime schedulers. Maybe we need something similar to crater at tokio
Because of batching I'd like to use an API that's heading towards compio but with a builder-plan style type io.
let mut io = Io::new();
let file = io.file("foo.txt");
let open = file.open(&mut io);
let read = file.read_to_string(&mut io, 64 * 1024);
let close = file.close(&mut io);
io.soft_link([open, read])
.hard_link([read, close]);
let result = io.submit().await?;
This API should take care of handling edge cases to prevent dead locking and handling behavior for soft link errors.
The bookkeeping ensures runtime waker does not get invalidated during the execution of the operation which can cause your future to never be polled again.
IO Uring Buf ring allows you to manage ring buffers to gain advantage with sockets See here
I'm visioning O_DIRECT reads to look like
let file = io.file("huge.dat")
.read_only()
.direct();
let mut pool = FixedPool::new()
.entries(64)
.buffer_size(256 * 1024);
let open = io.open(file);
let read = io.read_fixed(
file,
pool.buffer(3),
4 * 1024 * 1024,
256 * 1024,
);
let close = io.close(file);
io.soft_link([open, read])
.hard_link([read, close]);
let (result, pool) = io.submit(pool).await?;
let buf = pool.buffer(3);
println!("{:?}", &buf[..result.bytes(read)?]);
So is everyone just doing file I/O with io-uring what about sockets? there are some nice optimizations that come with sockets but first we should finalize on an API to support completion based I/O with a work stealing battle tested runtime on tokio_unstable. More on sockets soon
Links: