
Jana Toljaga, Nicolas Derumigny, Tara Aggoun, Mathieu Bacou, Gaël Thomas
The 32nd Symposium on Operating Systems Principles (SOSP 2026) 2026
A crash-consistent I/O cache is essential to ensure data integrity while optimizing performance. However, a general-purpose kernel fails to ensure integrity efficiently because of the high cost of synchronizing I/O requests with the many subsystems that rely on the cache. As a consequence, applications that must store data consistently mostly bypass the kernel I/O cache to implement a second cache in user space. This approach defeats the purpose of centralizing system mechanisms within the OS, induces overheads due to data copying and serialization, and leads to high engineering costs as user-space processes lack direct access to hardware features like page tables. In this paper, we address these drawbacks by doubling the number of kernels instead of doubling the number of caches. The first kernel is general-purpose, while the second provides an independent I/O stack tailored to efficiently ensure integrity. We implemented this design in VoliStorM using hardware virtualization. Our evaluation shows that, across databases, in-memory cache checkpointing, and language runtimes, VoliStorM outperforms the dual-cache approach while simplifying application code and improving code reuse.
Jana Toljaga, Nicolas Derumigny, Tara Aggoun, Mathieu Bacou, Gaël Thomas
The 32nd Symposium on Operating Systems Principles (SOSP 2026) 2026
A crash-consistent I/O cache is essential to ensure data integrity while optimizing performance. However, a general-purpose kernel fails to ensure integrity efficiently because of the high cost of synchronizing I/O requests with the many subsystems that rely on the cache. As a consequence, applications that must store data consistently mostly bypass the kernel I/O cache to implement a second cache in user space. This approach defeats the purpose of centralizing system mechanisms within the OS, induces overheads due to data copying and serialization, and leads to high engineering costs as user-space processes lack direct access to hardware features like page tables. In this paper, we address these drawbacks by doubling the number of kernels instead of doubling the number of caches. The first kernel is general-purpose, while the second provides an independent I/O stack tailored to efficiently ensure integrity. We implemented this design in VoliStorM using hardware virtualization. Our evaluation shows that, across databases, in-memory cache checkpointing, and language runtimes, VoliStorM outperforms the dual-cache approach while simplifying application code and improving code reuse.

Subashiny Tanigassalame, Yohan Pipereau, Adam Chader, Jana Toljaga, Gaël Thomas
Proceedings of the 25th International Middleware Conference (Middleware 2024) 2024
Partitioning a multi-threaded application between a secure and a non-secure memory zone remains a challenge. The current tools rely on data flow analysis techniques, which are unable to handle multi-threaded C or C++ applications. To avoid this limitation, we propose to trade the ease-of-use of data flow analysis for another language construct: explicit secure typing. With secure typing, as with data flow analysis, the developer annotates memory locations that contain sensitive values. However, instead of analyzing how the sensitive values flow, we propose to use these annotations to only check typing rules, such as ensuring that the code never stores a sensitive value in an unsafe memory location. By avoiding data flow analysis, the developer has to annotate more memory locations, but the partitioning tool can handle multi-threaded C and C++ applications.
Subashiny Tanigassalame, Yohan Pipereau, Adam Chader, Jana Toljaga, Gaël Thomas
Proceedings of the 25th International Middleware Conference (Middleware 2024) 2024
Partitioning a multi-threaded application between a secure and a non-secure memory zone remains a challenge. The current tools rely on data flow analysis techniques, which are unable to handle multi-threaded C or C++ applications. To avoid this limitation, we propose to trade the ease-of-use of data flow analysis for another language construct: explicit secure typing. With secure typing, as with data flow analysis, the developer annotates memory locations that contain sensitive values. However, instead of analyzing how the sensitive values flow, we propose to use these annotations to only check typing rules, such as ensuring that the code never stores a sensitive value in an unsafe memory location. By avoiding data flow analysis, the developer has to annotate more memory locations, but the partitioning tool can handle multi-threaded C and C++ applications.

Subashiny Tanigassalame, Yohan Pipereau, Adam Chader, Jana Toljaga, Gaël Thomas
International Conference on Advanced Information Networking and Applications 2024
Designing an efficient privacy-preserving application with Intel SGX is difficult. The problem comes from the prohibitive cost of switching the processor from the non-secure mode to the secure mode. To avoid this cost, we propose to design an SGX application as a distributed system with worker threads that communicate by exchanging messages. We implemented FastSGX, a runtime that exposes this programming model to the developer, and evaluated it with several data structures. Our evaluation with different workloads shows that the applications designed with FastSGX consistently outperform, and by up to 2.8x, the equivalent applications designed with the software development kit provided by Intel to use SGX.
Subashiny Tanigassalame, Yohan Pipereau, Adam Chader, Jana Toljaga, Gaël Thomas
International Conference on Advanced Information Networking and Applications 2024
Designing an efficient privacy-preserving application with Intel SGX is difficult. The problem comes from the prohibitive cost of switching the processor from the non-secure mode to the secure mode. To avoid this cost, we propose to design an SGX application as a distributed system with worker threads that communicate by exchanging messages. We implemented FastSGX, a runtime that exposes this programming model to the developer, and evaluated it with several data structures. Our evaluation with different workloads shows that the applications designed with FastSGX consistently outperform, and by up to 2.8x, the equivalent applications designed with the software development kit provided by Intel to use SGX.