2.3 KiBLFS
2.3 KiBLFS
Data Download in Modal
Why Download Inside Functions?
Modal functions run in isolated containers. Data must be downloaded inside the function or stored in Modal volumes.
HuggingFace Hub Download
Use huggingface_hub to download datasets:
# Add huggingface_hub to image
image = modal.Image.debian_slim(python_version="3.11").pip_install(
"torch",
"einops",
"numpy",
"huggingface_hub",
)
@app.function(gpu="A100", image=image, timeout=3600)
def train():
import os
from huggingface_hub import hf_hub_download
# Download tokenized dataset shards (publicly accessible, no auth required)
data_dir = "/tmp/data/dataset"
os.makedirs(data_dir, exist_ok=True)
def get_file(fname):
if not os.path.exists(os.path.join(data_dir, fname)):
print(f"Downloading {fname}...")
hf_hub_download(
repo_id="your-org/your-dataset",
filename=fname,
repo_type="dataset",
local_dir=data_dir,
)
# Download validation shard
get_file("val_000000.bin")
# Download first training shard
get_file("train_000001.bin")
# Load and use the data from data_dir
...
Example Dataset Layout
Tokenized datasets often ship in shard files like:
| File | Purpose |
|---|---|
val_000000.bin |
Validation shard |
train_000001.bin |
Training shard |
Private Datasets
For private HuggingFace datasets:
@app.function(gpu="A100", image=image, timeout=3600, secrets=[modal.Secret.from_name("huggingface")])
def train():
import os
from huggingface_hub import hf_hub_download
# Token is automatically available from secret
hf_hub_download(
repo_id="your-private-repo",
filename="data.bin",
repo_type="dataset",
local_dir="/tmp/data",
token=os.environ.get("HF_TOKEN"),
)
Modal Volumes (Persistent Storage)
For large datasets you want to cache:
volume = modal.Volume.from_name("training-data", create_if_missing=True)
@app.function(gpu="A100", image=image, volumes={"/data": volume})
def train():
# Data persists across function calls
if not os.path.exists("/data/train_000001.bin"):
download_data("/data")
# Use cached data
...