Dataset Hub
Datasets share the repository model with models — a namespace, a name, files and a card — plus a viewer that reads the actual rows.
Creating a dataset
curl -X POST https://inferix.co/api/v0/datasets \
-H "Authorization: Bearer $INFERIX_TOKEN" \
-H "Content-Type: application/json" \
-d '{"name": "my-dataset", "visibility": "PRIVATE"}'As with models, a verified email is required to publish publicly. Private datasets are not gated on it.
Supported formats
Parquet, CSV, JSON and JSONL are understood by the viewer. Parquet is the one to reach for on anything large — the viewer pages through it without downloading the file, so row previews and column statistics stay fast at any size.
A conventional layout keeps splits discoverable:
data/train-00000-of-00002.parquet data/train-00001-of-00002.parquet data/validation-00000-of-00001.parquet README.md
Uploading
Same pre-signed flow as models — the bytes never pass through the API:
curl -X POST https://inferix.co/api/v0/datasets/$NS/$NAME/files/upload-url \
-H "Authorization: Bearer $INFERIX_TOKEN" \
-H "Content-Type: application/json" \
-d '{"path": "data/train-00000-of-00001.parquet", "size": 1048576}'Git and Git LFS work as well, at https://inferix.co/api/git/datasets/<ns>/<name>.git.
Data Studio
Every dataset page opens on a viewer that reads the real files:
- Paged row preview, per split and per config.
- Column statistics — distributions for categorical columns, min/max/mean and histograms for numeric ones, and the proportion of nulls.
- A SQL console that runs in your browser against the Parquet files.
Row counts come from the dataset card. A dataset whose card does not declare one shows no count rather than zero.
Dataset cards
---
license: cc-by-4.0
task_categories:
- question-answering
language:
- en
dataset_info:
splits:
- name: train
num_examples: 87599
- name: validation
num_examples: 10570
---
# My dataset
Where it came from, how it was collected, and what it may be used for.dataset_info.splits[].num_examples is what populates the row count shown on cards and search results.
Visibility and gating
Public, private and gated behave exactly as they do for models, and cover the sub-resources too — files, commits, discussions and pull requests all follow the repository's visibility.
Gated datasets are the usual choice for anything with a licence that requires accepting terms: the listing stays discoverable, the files need an approved request.
Next
Model Hub — publish what you trained.
Fine-tuning — train a LoRA adapter on a dataset you have uploaded.
REST API reference — every endpoint.