Jump to content

Main menu Navigation ●Main page ●Contents ●Current events ●Random article ●About Wikipedia ●Contact us ●Donate Contribute ●Help ●Learn to edit ●Community portal ●Recent changes ●Upload file

●Create account ●Log in ●Create account ● Log in Pages for logged out editors learn more ●Contributions ●Talk

(Top) 1 Design 2 Interface 3 Performance 4 See also 5 References 6 Bibliography 7 External links

Google File System

●Български ●Català ●Čeština ●Deutsch ●Español ●Français ●한국어 ●हिन्दी ●Italiano ●עברית ●日本語 ●Português ●Русский ●Українська ●中文 Edit links ●Article ●Talk ●Read ●Edit ●View history Tools Actions ●Read ●Edit ●View history General ●What links here ●Related changes ●Upload file ●Special pages ●Permanent link ●Page information ●Cite this page ●Get shortened URL ●Download QR code ●Wikidata item Print/export ●Download as PDF ●Printable version In other projects ●Wikimedia Commons Appearance From Wikipedia, the free encyclopedia

This article needs additional citations for verification. Please help improve this articlebyadding citations to reliable sources. Unsourced material may be challenged and removed.
Find sources: "Google File System" – news · newspapers · books · scholar · JSTOR (July 2016) (Learn how and when to remove this message)

Operating system	Linux kernel
Type	Distributed file system
License	Proprietary

Google File System (GFSorGoogleFS, not to be confused with the GFS Linux file system) is a proprietary distributed file system developed by Google to provide efficient, reliable access to data using large clusters of commodity hardware. Google file system was replaced by Colossus in 2010.^[1]

Design

[edit]

Google File System is designed for system-to-system interaction, and not for user-to-system interaction. The chunk servers replicate the data automatically.

GFS is enhanced for Google's core data storage and usage needs (primarily the search engine), which can generate enormous amounts of data that must be retained; Google File System grew out of an earlier Google effort, "BigFiles", developed by Larry Page and Sergey Brin in the early days of Google, while it was still located in Stanford. Files are divided into fixed-size chunks of 64 megabytes, similar to clusters or sectors in regular file systems, which are only extremely rarely overwritten, or shrunk; files are usually appended to or read. It is also designed and optimized to run on Google's computing clusters, dense nodes which consist of cheap "commodity" computers, which means precautions must be taken against the high failure rate of individual nodes and the subsequent data loss. Other design decisions select for high data throughputs, even when it comes at the cost of latency.

A GFS cluster consists of multiple nodes. These nodes are divided into two types: one Master node and multiple Chunkservers. Each file is divided into fixed-size chunks. Chunkservers store these chunks. Each chunk is assigned a globally unique 64-bit label by the master node at the time of creation, and logical mappings of files to constituent chunks are maintained. Each chunk is replicated several times throughout the network. At default, it is replicated three times, but this is configurable.^[2] Files which are in high demand may have a higher replication factor, while files for which the application client uses strict storage optimizations may be replicated less than three times - in order to cope with quick garbage cleaning policies.^[2]

The Master server does not usually store the actual chunks, but rather all the metadata associated with the chunks, such as the tables mapping the 64-bit labels to chunk locations and the files they make up (mapping from files to chunks), the locations of the copies of the chunks, what processes are reading or writing to a particular chunk, or taking a "snapshot" of the chunk pursuant to replicate it (usually at the instigation of the Master server, when, due to node failures, the number of copies of a chunk has fallen beneath the set number). All this metadata is kept current by the Master server periodically receiving updates from each chunk server ("Heart-beat messages").

Permissions for modifications are handled by a system of time-limited, expiring "leases", where the Master server grants permission to a process for a finite period of time during which no other process will be granted permission by the Master server to modify the chunk. The modifying chunkserver, which is always the primary chunk holder, then propagates the changes to the chunkservers with the backup copies. The changes are not saved until all chunkservers acknowledge, thus guaranteeing the completion and atomicity of the operation.

Programs access the chunks by first querying the Master server for the locations of the desired chunks; if the chunks are not being operated on (i.e. no outstanding leases exist), the Master replies with the locations, and the program then contacts and receives the data from the chunkserver directly (similar to Kazaa and its supernodes).

Unlike most other file systems, GFS is not implemented in the kernel of an operating system, but is instead provided as a userspace library.^[3]

Interface

[edit]

The Google File System does not provide a POSIX interface.^[4] Files are organized hierarchically in directories and identified by pathnames. The file operations such as create, delete, open, close, read, write are supported. It supports Record Append which allows multiple clients to append data to the same file concurrently and atomicity is guaranteed.

Performance

[edit]

Deciding from benchmarking results,^[2] when used with relatively small number of servers (15), the file system achieves reading performance comparable to that of a single disk (80–100 MB/s), but has a reduced write performance (30 MB/s), and is relatively slow (5 MB/s) in appending data to existing files. The authors present no results on random seek time. As the master node is not directly involved in data reading (the data are passed from the chunk server directly to the reading client), the read rate increases significantly with the number of chunk servers, achieving 583 MB/s for 342 nodes. Aggregating multiple servers also allows big capacity, while it is somewhat reduced by storing data in three independent locations (to provide redundancy).

References

[edit]

^ Ma, Eric (2012-11-29). "Colossus: Successor to the Google File System (GFS)". SysTutorials. Archived from the original on 2019-04-12. Retrieved 2016-05-10.

^ ^a ^b ^c Ghemawat, Gobioff & Leung 2003.

^ Kyriazis, Dimosthenis (2013). Data Intensive Storage Services for Cloud Environments. IGI Global. p. 13. ISBN 9781466639355.

^ Marshall Kirk McKusick; Sean Quinlan (August 2009). "GFS: Evolution on Fast-forward". ACM Queue. 7 (7): 10–20. doi:10.1145/1594204.1594206. Retrieved 21 December 2019.

Bibliography

[edit]

Ghemawat, S.; Gobioff, H.; Leung, S. T. (2003). "The Google file system". Proceedings of the nineteenth ACM Symposium on Operating Systems Principles - SOSP '03 (PDF). p. 29. CiteSeerX 10.1.1.125.789. doi:10.1145/945445.945450. ISBN 1581137575. S2CID 221261373.

External links

[edit]

"GFS: Evolution on Fast-forward", Queue, ACM.
"Google File System Eval, Part I", Storage mojo.

File systems

Disk and
non-rotating

ADFS
AdvFS
Amiga FFS
Amiga OFS
APFS
AthFS
bcachefs
BFS
- Be File System
- Boot File System
- Byte File System (z/VM)
Btrfs
CVFS
CXFS
DFS
EFS
- Encrypting File System
- Extent File System
Episode
ext
- ext2
- ext3
- ext3cow
- ext4
FAT
- exFAT
Files-11
Fossil
GPFS
HAMMER
- HAMMER2
HFS (Classic Mac OS)
HFS (MVS)
HFS+
HPFS
HTFS
JFS
LFS
MFS
- Macintosh File System
- TiVo Media File System
MINIX
NetWare File System
Next3
NILFS
- NILFS2
NSS
NTFS
OneFS
OpenZFS
PFS
QFS
QNX4FS
ReFS
ReiserFS
- Reiser4
Reliance
Reliance Nitro
RFS
SFS
- Shared File System (VM)
- Smart File System
SNFS
Soup (Apple)
Tux3
UBIFS
UFS/UFS2
- soft updates
- WAPBL
VxFS
WAFL
Xiafs
XFS
Xsan
zFS (z/OS)
ZFS (Sun)

Optical disc

Flash memory and SSD

host-side wear leveling

Distributed parallel

NAS

Specialized

Pseudo

Encrypted

Types

Features

Case preservation
Copy-on-write
Data deduplication
Data scrubbing
Execute in place
Extent
File attribute
- Extended file attributes
File change log
Fork
Links
- Hard
- Symbolic

Layouts

Company

Divisions

Ads
AI
- Brain
- DeepMind
Android
China
- Goojje
Chrome
Cloud
Glass
Google.org
Health
Maps
Pixel
Search
- Timeline
Sidewalk Labs
Sustainability
YouTube
- History
- "Me at the zoo"
- Social impact
- YouTuber

People

Current

Former

Real estate

Design

Fonts
- Croscore
- Noto
- Product Sans
- Roboto
Logo
- Doodle
  - Doodle Champion Island Games
  - Magic Cat Academy
Material Design

Events

YouTube

Projects and
initiatives

20% project
Area 120
- Reply
- Tables
ATAP
Business Groups
Computing University Initiative
Data Liberation Front
Data Transfer Project
Developer Expert
Digital Garage
Digital News Initiative
Digital Unlocked
Dragonfly
Founders' Award
Free Zone
Get Your Business Online
Google for Education
Google for Startups
Labs
Liquid Galaxy
Made with Code
Māori
ML FairnessNative Client
News Lab
Nightingale
OKR
PowerMeter
Privacy Sandbox
Quantum Artificial Intelligence Lab
RechargeIT
Shield
Silicon Initiative
Solve for X
Starline
Student Ambassador Program
Submarine communications cables
- Dunant
- Grace Hopper
Sunroof
YouTube
Zero

Criticism

YouTube

Development

Operating systems

Android
- Automotive
- Glass OS
- Go
- gLinux
- Goobuntu
- Things
- TV
- Wear OS
ChromeOS
- ChromiumOS
- Neverware
Fuchsia
TV

Libraries/
frameworks

Platforms

App Engine
AppJet
Apps Script
Cloud Platform
- Anvato
Firebase
- Cloud Messaging
- Crashlytics
Global IP Solutions
- Internet Low Bitrate Codec
- Internet Speech Audio Codec
Gridcentric, Inc.
ITA Software
Kubernetes
LevelDB
Neatx
Project IDX
SageTV

Apigee

Tools

Search algorithms

Others

BERT
BigQuery
Chrome Experiments
Flutter
Gemini
Googlebot
Keyhole Markup Language
LaMDA
Open Location Code
PaLM
Programming languages
- Caja
- Carbon
- Dart
- Go
- Sawzall
Transformer
Viewdle
Webdriver Torso
Web Server

File formats

AAB
APK
- AV1
On2 Technologies
- VP3
- VP6
- VP8
  - libvpx
VP9
WebM
WebP
WOFF2

Products

Entertainment

Play

YouTube

Communication

Aardvark
Alerts
Answers
Base
BeatThatQuote.com
Blog Search
Books
- Ngram Viewer
Code Search
Data Commons
Dataset Search
Dictionary
Directory
Fast Flip
Flu Trends
Finance
Goggles
Google.by
Images
- Image Labeler
- Image Swirl
Kaltix
Knowledge Graph
- Freebase
- Metaweb
Like.com
News
- Archive
- Weather
Patents
People Cards
Personalized Search
Public Data Explorer
Questions and Answers
SafeSearch
Scholar
Searchwiki
Shopping
Catalogs
- Express
Squared
Tenor
Travel
- Flights
Trends
- Insights for Search
Voice Search
WDYL

Navigation

Earth
Endoxon
ImageAmerica
Maps
- Latitude
- Map Maker
- Navigation
- Pin
- Street View
  - Coverage
  - Trusted
Waze

Business
and finance

Ad Manager
AdMob
Ads
Adscape
AdSense
Attribution
BebaPay
Checkout
Contributor
DoubleClick
- Affiliate Network
- Invite Media
Marketing Platform
- Analytics
- Looker Studio
- Urchin
Pay (mobile app)
- Wallet
- Pay (payment method)
- Send
- Tez
PostRank
Primer
Softcard
Wildfire Interactive
Widevine

Organization
and productivity

Docs Editors

Publishing

Education

Others

Chrome

Images and
photography

Hardware

Smartphones	Android Dev Phone Android One Nexus Nexus One S Galaxy Nexus 4 5 6 5X 6P Comparison Pixel Pixel 2 3 3a 4 4a 5 5a 6 6a 7 7a Fold 8 8a Comparison Play Edition Project Ara
Laptops and tablets	Chromebook Nexus 7 (2012) 7 (2013) 10 9 Comparison Pixel Chromebook Pixel Pixelbook Pixelbook Go C Slate Tablet
Wearables	Fitbit List of products Pixel Buds Pixel Watch Pixel Watch 2 Project Iris (unreleased) Virtual reality Cardboard Contact Lens Daydream Glass
Others	Chromebit Chromebox Clips Digital media players Chromecast Nexus Player Nexus Q Dropcam Liquid Galaxy Nest Smart Speakers Thermostat Wifi OnHub Pixel Visual Core Search Appliance Sycamore processor Tensor Tensor Processing Unit Titan Security Key

v t e Litigation
Advertising	Feldman v. Google, Inc. (2007) Rescuecom Corp. v. Google Inc. (2009) Goddard v. Google, Inc. (2009) Rosetta Stone Ltd. v. Google, Inc. (2012) Google, Inc. v. American Blind & Wallpaper Factory, Inc. (2017) Jedi Blue
Antitrust	European Union (2010–present) United States v. Adobe Systems, Inc., Apple Inc., Google Inc., Intel Corporation, Intuit, Inc., and Pixar (2011) Umar Javeed, Sukarma Thapar, Aaqib Javeed vs. Google LLC and Ors. (2019) United States v. Google LLC (2020) United States v. Google LLC (2023)
Intellectual property	Perfect 10, Inc. v. Amazon.com, Inc. and A9.com Inc. and Google Inc. (2007) Viacom International Inc. v. YouTube, Inc. (2010) Lenz v. Universal Music Corp.(2015) Authors Guild, Inc. v. Google, Inc. (2015) Field v. Google, Inc. (2016) Google LLC v. Oracle America, Inc. (2021) Smartphone patent wars
Privacy	Rocky Mountain Bank v. Google, Inc. (2009) Hibnick v. Google, Inc. (2010) United States v. Google Inc. (2012) Judgement of the German Federal Court of Justice on Google's autocomplete function (2013) Joffe v. Google, Inc. (2013) Mosley v SARL Google (2013) Google Spain v AEPD and Mario Costeja González (2014) Frank v. Gaos (2019)
Other	Garcia v. Google, Inc. (2015) Google LLC v Defteros (2020) Epic Games v. Google (2021) Gonzalez v. Google LLC (2022)
Category

Terms and phrases	"Don't be evil" Gayglers Google (verb) Google bombing 2004 U.S. presidential election Google effect Googlefight Google hacking Googleshare Google tax Googlewhack Googlization "Illegal flower tribute" Rooting Search engine manipulation effect Sitelink Site reliability engineering YouTube poop
Documentaries	AlphaGo Google: Behind the Screen Google Maps Road Trip Google and the World Brain The Creepy Line
Books	Google Hacks The Google Story Google Volume One Googled: The End of the World as We Know It How Google Works I'm Feeling Lucky In the Plex The Google Book The MANIAC
Popular culture	Google Feud Google Me (film) "Google Me" (Kim Zolciak song) "Google Me" (Teyana Taylor song) Is Google Making Us Stupid? Proceratium google Matt Nathanson: Live at Google The Billion Dollar Code The Internship Where on Google Earth is Carmen Sandiego?
Others	"Attention Is All You Need" elgooG Predictions of the end Registry .app (top-level domain) .dev g.co .google Pimp My Search Relationship with Wikipedia Sensorvault Stanford Digital Library Project

Italics indicate discontinued products or services.
Category
Commons
Outline
WikiProject

Retrieved from "https://en.wikipedia.org/w/index.php?title=Google_File_System&oldid=1229741435" Categories: ●Distributed file systems supported by the Linux kernel ●Google ●Parallel computing ●Distributed file systems Hidden categories: ●Articles with short description ●Short description is different from Wikidata ●Articles needing additional references from July 2016 ●All articles needing additional references ●This page was last edited on 18 June 2024, at 13:53 (UTC). ●Text is available under the Creative Commons Attribution-ShareAlike License 4.0; additional terms may apply. By using this site, you agree to the Terms of Use and Privacy Policy. Wikipedia® is a registered trademark of the Wikimedia Foundation, Inc., a non-profit organization. ●Privacy policy ●About Wikipedia ●Disclaimers ●Contact Wikipedia ●Code of Conduct ●Developers ●Statistics ●Cookie statement ●Mobile view