Researchers at Stanford who need to understand where their data should live will find Stanford Research Computing Data Storage a more useful starting point than the typical institutional storage page manages to be. Stanford Research Computing Data Storage covers three distinct tiers run by Stanford Research Computing (SRCC) for Principal Investigators, their groups, and affiliated researchers whose work feeds into active computation or long-term retention obligations. It is built around the assumption that you already have a compute allocation and a research problem producing far more data than any departmental file share could hold.

Sherlock scratch storage

The first tier, Sherlock scratch storage, is the most short-lived of the three. It attaches to the Sherlock HPC cluster and gives each user and PI group a pair of working directories, $SCRATCH and $GROUP_SCRATCH, for the files and code tied to running jobs. The retention limit is built in by design: data sitting untouched for 90 days can be cleared, and Stanford Research Computing Data Storage is direct about the fact that this is not a backup target. That honesty earns attention. Plenty of people misuse scratch space as a safe harbour and then lose work, so spelling out the limit up front heads off a predictable category of pain.

Oak persistent tier

Oak is where Stanford Research Computing Data Storage carries the most technical depth. It is the persistent tier, plugged into both the Sherlock and SCG compute clusters, and the documentation is unusually candid that SRCC built it in-house on Lustre with the Robinhood Policy Engine. Naming the underlying open-source components is a deliberate choice: it tells you the intended audience is technical enough to care what they are running on. Oak is pitched at large datasets staged for active compute, curated results after processing, and the outputs that sit underneath published papers. Access methods are broad: Globus, SFTP, RSYNC, SCP, SSHFS, and RCLONE all work out of the box, with SMB and NFSv4 gateways available at an additional fee.

Monthly billing and data classification limits

Two details on Oak in Stanford Research Computing Data Storage affect real planning decisions. It is billed monthly, so this is an operational cost a research group has to budget for, not a free university perk. And it is approved only for Low and Moderate Risk Data under Stanford's own Information Security Office classifications. Anyone handling more sensitive material needs to read that line carefully, because it draws a hard boundary on what the tier will accept. The page does not bury this, which is the right call. That risk-classification note sits plainly in the description instead of being tucked into a footnote, and it changes how a lab with restricted data should read the whole page.

Elm archival storage

Elm rounds out Stanford Research Computing Data Storage as the cold, archival option, and the numbers attached to it are striking. It is designed for datasets running from terabytes up into hundreds of petabytes, written to tape in Stanford's own data center. The design assumption is write-once, read-rarely: ingest a large body of data, then leave it alone except when compliance or policy requires a retrieval. High-throughput Globus transfers handle both directions. The argument against cloud archival is concrete and free of marketing language: Elm has no egress or restore fees beyond the monthly allocation cost, which is exactly what makes cloud cold storage deceptively expensive once you need your data back. For a lab sitting on regulatory retention obligations, that fee structure is the entire case for choosing tape.

Choosing between Oak or Elm

What makes Stanford Research Computing Data Storage genuinely useful as a reference, beyond the individual tier descriptions, is that it tries to help you choose. There is a decision guide labelled "Oak or Elm? Help me choose," which addresses precisely the question a researcher hits when they have data too valuable to delete but too large to keep hot. The three tiers are not interchangeable, and the cost and access tradeoffs between them require thought, so a guide that pushes people toward the right tier instead of leaving them to guess does real work. A companion section documents the additional tools for moving and managing data, so Stanford Research Computing Data Storage is presented as part of a workflow and not three isolated buckets.

Pricing details live elsewhere

If you arrived at this page through a business directory or a general web search expecting a self-contained pricing sheet, you will leave with a gap. Stanford Research Computing Data Storage explains what each tier is and what it covers, but the figures that drive a budget decision, such as quotas, per-terabyte rates, and allocation minimums, live elsewhere or behind an account. Someone arriving cold will understand the shape of the offering and the philosophy behind it, then need to go further to get numbers. For an internal university service, that is entirely normal, since the audience is assumed to already be inside the Stanford system. It is an orientation document more than a complete specification, and reading it that way is the right frame.

How the three tiers fit together

There is a coherence to the three tiers that becomes clear once you read them together. Scratch handles the volatile working set, Oak holds the data you are actively analysing alongside results worth keeping accessible, and Elm absorbs the long tail that has to survive for years but rarely gets touched. That progression from fast and temporary to slow and permanent maps onto the standard research data lifecycle without forcing every dataset into one bucket. The in-house engineering on Oak, with the choice to document the underlying file systems, points to a team that has thought carefully about how research storage should work instead of simply repackaging a vendor offering.

Documentation quality for Stanford researchers

Stanford Research Computing Data Storage earns a favourable reading, though the scope is narrow and the narrowness is worth naming plainly. For Stanford PIs and their groups, it is the right reference to start from when deciding where research data should live. The decision guide, the candid notes on retention limits, the risk-classification boundary, and the Elm fee argument against cloud egress all make Stanford Research Computing Data Storage more trustworthy than a glossier page would be. The plain warning that scratch storage is not a backup, repeated clearly instead of buried, reads like something written by people who have watched researchers learn it the hard way. For institutions building or reviewing their own tiered storage documentation, Stanford Research Computing Data Storage is a useful model of how to explain layered infrastructure without overselling any part of it.