2026 年 34 巻 p. 14-22
Distributed file systems employ two deduplication methods: intra-node deduplication and inter-node deduplication. Inter-node deduplication, which eliminates redundant data across multiple nodes, achieves a superior capacity reduction compared to intra-node deduplication, which is confined to a single node. However, inter-node deduplication has the risk to store duplicate data on a different node from the original file, which increases the number of inter-node communications and reduces the data read performance. To mitigate this challenge, we propose an approach that retains duplicate data within the original file until storage capacity constraints necessitate eviction, thereby minimizing inter-node communication. The evaluation results show that the proposed method improves I/O performance by 3.2 times compared to the conventional method in the training workload of AI image processing. Furthermore, our evaluation results indicate that the release of the cache in units of duplicate chunks may result in a decline in data read performance when compared to the conventional method. However, this issue can be addressed by releasing the cache in units of files.