Distributed File Systems

Authored by,
Srishti Yadav (Student, Computer Department, Vidyalankar Institute of Technology)
Vedanti Ghanekar (Student, Computer Department, Vidyalankar Institute of Technology)
Isha Patil (Student, Computer Department, Vidyalankar Institute of Technology)
Pradnya Dhengale (Student, Computer Department, Vidyalankar Institute of Technology)
Under guidance of,
Dr. Prakash Parmar (Faculty, Computer Department, Vidyalankar Institute of Technology)
Dr. Amit K. Nerurkar (Faculty, Computer Department, Vidyalankar Institute of Technology)
What is Distributed File System?
A distributed system is a network of multiple components that interact with each other to appear as a single system to user. In today's AI driven world, storing and accessing massive amounts of information efficiently and securely is more important than ever. Distributed file systems (DFS) make this possible by spreading data across multiple machines while ensuring reliability, fault tolerance, and high performance.
Goals of Distributed File System
Transparency: Users are unaware of where the data is physically located or what is the name of file providing single view of files.
Scalability: System will be able to work even after more computers are added without slowing down the performance or increasing workloads.
Reliability: Data is secure even when the system goes down.
Security: Security helps to keep data secure by giving its access to right people over the network.
Real world examples
Network File System (NFS): NFS was initially developed by a group of engineers at Sun Microsystems to provide file sharing capabilities for the company's UNIX systems in 1884. NFS helps multiple employees to access shared project files, documents, and resources from any workstation.
Hadoop Distributed File System: HDFS is a core component of the Apache Hadoop project, developed by the Apache Software Foundation in 2006. HDFS has master/slave architecture. HDFS supports parallel processing and has metadata level locking mechanism. Communication protocol (Hadoop RPC protocol) is mainly based on a custom Remote Procedure Call (RPC) protocol over TCP/IP.
Google File System: It was developed by Google, first publicly disclosed in 2003. A GFS cluster consists of a master node which maintains metadata (file namespace, mapping files to chunks, locations of replicas) and multiple chunk servers. GFS simplifies large-scale data processing.
Future Trends
AI and machine learning integration – Distributed file systems are integrating with artificial intelligence and machine learning pipelines helping it to become easier to analyze vast amounts of unstructured data.
Quantum computing and network – It may promise innovative approaches in data security, radical security during transmission for future networks.
Open source and community innovation – Open source DFS are booming which allows organizations to adapt systems for their needs and benefits, this leads to broader innovation.
References
Tanenbaum, A.S., & Van Steen, M. (2017). Distributed Systems: Principles and Paradigms (3rd ed.). distributed-systems.net
Pan, X., Luo, Z., & Zhou, L. (2024). Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations. Innovation and Application of Engineering and Technology, 2(1), 1–12. https://doi.org/10.62836/iaet.v2i1.157