> ## Documentation Index
> Fetch the complete documentation index at: https://starforge.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Roles and quotas

> Who may do what, and how much of the cluster they may hold.

```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
sf admin user-add alice
sf admin set-role alice operator
sf admin set-quota alice --max-concurrent-gpus 8 --max-concurrent-jobs 4
```

## Roles

Two of them.

| Role       | Can                                                                                                                              |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------- |
| `operator` | Submit, watch, compare, publish plugins and environments, manage their own datasets and models. Shown as "Member" in the console |
| `admin`    | All of that, plus users, quotas, hardware, settings, scheduled tasks, audit, and disabling published extensions                  |

An admin-only endpoint returns `403` to an `operator` token, not `401` — the token is fine, the
permission is not.

`operator` is the default for a new account, and `FORGE_OIDC_DEFAULT_ROLE` decides what a
first-time single sign-on login gets.

## Quota dimensions

A quota is per user, and five things are limited independently:

<ParamField path="max_concurrent_gpus" type="int" default="8">
  The most GPUs this user may hold at once, across all their running jobs. Exact — admission counts
  it before dispatch.
</ParamField>

<ParamField path="max_concurrent_jobs" type="int" default="4">
  How many jobs may run at once, regardless of size. Stops one person filling the queue with small
  jobs.
</ParamField>

<ParamField path="daily_gpu_hours" type="int" default="0">
  GPU-hours per day. 0 is unlimited. Day boundaries are computed in the service timezone.
</ParamField>

<ParamField path="max_storage_gb" type="int" default="0">
  How much the user's run directories may hold. 0 is unlimited. Measured rather than reserved, so it
  is eventually consistent.
</ParamField>

<ParamField path="allowed_profiles" type="comma-separated" default="all">
  Which hardware profiles this user may request. Empty means all of them.
</ParamField>

Plus `priority`, which orders the queue under `priority-fifo`, and a per-series map for a
heterogeneous cluster — `{series: {max_concurrent_gpus, daily_gpu_hours}}` — when one global number
is too blunt.

## Team quotas

Personal quotas are not enough once several departments share a cluster: eight
people at four cards each are all inside their own limits and are thirty-two
cards of one department's capacity. That is what a team quota governs.

**A run's budget comes from the team owning its project, never from whoever
submitted it.** This is the crux of the model: a person can belong to several
teams, and only the project answers whose capacity a run spent. The team is
stamped on the job row at submit, so moving a project to another department
later does not rewrite work already done -- the same rule `gpu_seconds` follows
by freezing on a terminal state.

Create teams, set their quota and manage members under **Administration →
Teams**; a project's owning team is changed in its settings, by an
administrator.

Two semantics are the opposite of the per-user ones, and both are worth
remembering:

* **No team quota means no team limit**, not a limit of zero. The per-user
  reading is the reverse (no quota, no submitting), which is a deliberate
  deny-by-default for new accounts; carried up here it would make creating a
  team a way to block everyone in it instantly.
* **Administrators are not exempt from a team quota.** A personal exemption
  says "this person is trusted with the cluster"; skipping a team's budget
  spends another department's money, and being an administrator is not a claim
  on that.

A job over its team's limit **queues** rather than being rejected: it is a
shared condition that clears as colleagues' jobs finish, and rejecting would
mean resubmitting by hand for something the platform can wait out itself.

## Billing

The platform has always measured GPU-seconds exactly and attributed them to a
project, a person and a team. Billing adds **one thing: a price.**

It re-measures nothing. The hours come from the same ledger the usage page
shows, and the attribution rule is `cost_breakdown`'s: **a run is counted in the
period it ended.** A long run spanning months belongs to the month it finished
in; a job still going is counted for the part it has run and labelled
separately. An invoice that disagrees with the usage page is worse than no
invoice -- it is an argument nobody can settle, and the usage page is the number
people already trust.

Read the bill under **Administration → Billing**, grouped by team, project or
person, and exportable as .xlsx.

### The price book

Rates are per **accelerator series** and carry a start date. Two rules:

* **The first rate for a card type covers all history** -- nothing was ever
  billed at another one, so there is no settled period to rewrite. That is what
  lets a deployment turning billing on price the quarter it already ran, which
  is the first thing anybody asks for.
* **A later rate cannot move backwards.** It would rewrite a bill somebody has
  already been shown and acted on. A change is a new row with a later start;
  rates accumulate and are never edited.

A card type that produced hours with no rate in force is **named**
(`unpriced_series`) and counts as zero. A confident total quietly missing a
series nobody remembered to price is much worse than one that says so.

Money is **integer micros** (millionths of a currency unit) from the ledger to
the report, rounded once at the edge. Floating point is not a thing to add up
money with: a fraction of a cent of drift per run becomes a total nobody can
reconcile.

### Monthly budgets

A team can carry a budget **per calendar month** in the service timezone. Over
it, new work **queues** rather than being rejected -- the month rolls over on
its own, and somebody told to resubmit in eleven days will not.

The reading is **already spent**, not predicted: nobody knows what a job that
has not run will cost. So the run that crossed the line always finishes -- the
overage is something it produced, not something it was admitted with. The same
shape as the daily GPU-hour budget, one level up and denominated in money.

## Borrowing and reclaim (preemption)

Static quotas fragment a cluster and then leave it idle -- the complaint every
shared cluster eventually produces, and the reason Kueue
(`reclaimWithinCohort`), Volcano (`deserved` plus reclaim), YARN (guaranteed vs
maximum capacity) and Slurm (QOS preemption) all grew **borrowing**.

With `FORGE_SCHED_PREEMPTION` on, a team's concurrent-GPU quota becomes a
**guarantee** rather than a ceiling:

* a team may run above it while the cluster has room (`capacity_gate` still
  bounds that by physics)
* that overflow is reclaimed when a team **below its own guarantee** is queued
  and blocked on capacity

**Borrowing and reclaim are one switch.** Lending without reclaim lets one team
hold the cluster indefinitely while others sit inside their quotas waiting,
which is worse than the hard cap it replaced.

### What is reclaimable, stated as narrowly as possible

All five must hold. Each exists because without it the mechanism does something
somebody would be right to be angry about:

| Condition                               | Without it                                                                                                  |
| --------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| the taker is under its guarantee        | this is not reclaim, it is one team outbidding another -- and no team could rely on the number it was given |
| the victim's team is over its guarantee | taking from a team inside its guarantee breaks the only promise a quota makes                               |
| a different team                        | a team preempting itself is a queue reordering itself, at the cost of a checkpoint                          |
| it has run for the minimum              | a busy queue would reclaim, admit, reclaim, and the cluster would spend its day writing checkpoints         |
| it holds the series that is short       | reclaiming an H200 job does not help somebody waiting for an H100                                           |

It takes the **youngest** first: the least work thrown away. Deliberately not
the lowest priority -- priority orders who runs next, and reusing it here would
let a high-priority long job be safe while a high-priority job that started a
minute ago is taken, answering two questions with one number.

### Reclaim is a pause

The platform already had this path and it already means the right thing:
`PAUSED(auto_resume)` is the *system-side* pause that maintenance drain and
window enforcement use. Metering stops at the pause, the working directory is
kept, and the queue worker resumes from the last checkpoint. **The person whose
job was reclaimed does nothing**, and is told why.

The cost is honest and bounded: progress since the last checkpoint. NeMo-RL
saves every `save_period` steps by default, so it is minutes.

## Splitting a GPU (MIG, time-slicing)

**It is one row in the hardware registry.** The NVIDIA GPU Operator advertises a
MIG partition as its own extended resource (`nvidia.com/mig-3g.40gb`) and a
time-slice as `nvidia.com/gpu.shared`. Register the partition as a card type
under **Administration → Settings → Hardware** with that name in
`k8s_resource`, and it gets independent capacity, quota, metering and time
windows -- because all four are already keyed by series id.

**The scheduler does not change at all.** Splitting a card is a registry entry,
not a scheduling feature.

Only the kuberay backend reads the field. On `local` and `node` the allocator
hands out whole cards by index and has no notion of a partition; declaring one
there would advertise capacity nothing can place.

## Approvals: an escalation path, not a workflow engine

A department that spends its month's budget on the third is queued for
twenty-eight days. Queueing is the right **default** -- the condition clears on
its own, and rejecting would mean telling somebody to resubmit in three weeks --
but there has to be a way to say "this one genuinely cannot wait", and for that
to leave a record of who said so.

**It is not wired into admission.** The obvious design is a third verb beside
`queue` and `reject`: over budget, so route to approval. It is wrong, and the
reason generalises. Team quota often clears within minutes as colleagues' jobs
finish; sending every such denial to a human produces a queue of requests that
answer themselves before anybody reads them, and an approver who learns to
ignore the list.

So escalation is **opt-in**: the job queues as it always did, and the submitter
asks for an exception from the job page when waiting is not an option.

**One approval lets one run past.** Not the team, not the month, and not until
somebody remembers to revoke it. A grant that outlives the reason it was given
is how a budget stops being a budget, and a narrow grant is one an approver can
actually reason about in the ten seconds they will spend on it.

The approver is an **owner of the team that holds the budget**, with
administrators as a fallback -- somebody has to be able to answer when a team has
no owner, or when the only owner is the person asking. **Nobody approves their
own request**; rejecting your own is allowed, because withdrawing needs no second
person.

A quota change is always an administrator's call: a team owner approving their
own raise is not a limit.

### Deliberately absent

One level, and an expiry. No counter-signing, chains, or delegation. Those
belong to whatever system an organisation already runs its approvals in;
rebuilding them here produces a worse copy that nobody's compliance department
accepts anyway. What the platform owes such a system is **a record and a
webhook**, not a re-implementation.

Requests nobody answers expire on their own. A pending list that only grows is
one nobody reads, and the request that mattered is somewhere in it.

## Syncing teams from single sign-on

With `FORGE_OIDC_GROUP_CLAIM` set, every sign-in reconciles the person's
membership against the directory -- **removing as well as adding**: somebody who
left a department in the IdP has left it here. Empty (the default) means the
platform never touches membership, which a deployment that manages teams by hand
must keep, since reconciling against an absent claim would empty everybody out.

A group the directory names but the platform does not have is skipped, never
created: an IdP group is not permission to invent a budget holder.
`FORGE_OIDC_GROUP_PREFIX` strips the path prefix a directory hands out.

## Soft by default

`FORGE_QUOTA_ENFORCE=false` records usage without refusing anything. That is useful while you work
out what the numbers should be, and dangerous to leave on afterwards.

Even with enforcement on, bare-metal Ray has no cgroup or GPU hard isolation, so admission is a gate
applied *before* the job starts. The [runtime watchdog](/en/ops/observability) is the second layer:
it reconciles what the cluster is actually using against what the ledger says, and can stop a job
that reliably exceeds its allocation. A hard quota needs infrastructure like Kubernetes or Kueue.

## Reading a refusal

| Message names       | Meaning                                                                |
| ------------------- | ---------------------------------------------------------------------- |
| Concurrent GPUs     | The user is at their ceiling. Not a cluster problem                    |
| Storage quota       | Run directories over budget. `sf clean` an exported experiment         |
| Daily GPU-hours     | Today's budget is spent; it resets at midnight in the service timezone |
| Profile not allowed | `allowed_profiles` does not include the requested one                  |
| Cluster capacity    | Genuinely no free cards. This one queues rather than refusing          |

The last row is the important distinction: **waiting in the queue is not a quota failure**. A queued
job was admitted.

## Audit

Creating a user, setting a role, enabling or disabling an account and setting a quota are all
recorded in the [audit log](/en/console/audit) with the administrator who did it. So is every
submission that was refused on quota — which is what makes "why did my job not run" answerable after
the error message has scrolled away.
