Methodology

How the league works

A season-long, 6-team PPR league where every manager is either a decision model or a fixed strategy. Nobody gets human help after the draft.

The managers

Same league context. Different decision rules.

ManagerWhat it isRuns on
glifanGLiNER2.5-Decide, fine-tuned on 14,481 real NFL start/sit outcomes. 340M parameters.Fastino API
JevTypeSafe's hosted System One decision model, zero-shot. Parameter count undisclosed.Vercel AI Gateway
Decide (untrained)The same GLiNER2.5-Decide model as glifan, with no fantasy training. The control group.Fastino API
FantasyPros expertsStarts whoever the FantasyPros weekly expert consensus projects to score the most.Human consensus
Season averageStarts whoever has averaged the most PPR points this season, skipping anyone ruled out.Rule
Hot handStarts whoever scored the most over the last three games, skipping anyone ruled out.Rule

Keeping it fair

  • Same draft rules. All six teams snake-draft from FantasyPros redraft consensus in a seeded random order, so rosters are comparable and the contest is weekly lineup decisions.
  • Same information. The AI managers see the identical pre-game description of each player and answer the identical question: will he finish as a start, a flex, or a bench player at his position this week?
  • Locked at kickoff. Lineups refresh with injury news until each game starts, then freeze. Every lineup is committed to GitHub before kickoff.
  • Standard scoring. PPR, one QB, two RB, two WR, one TE and one RB/WR/TE flex, head-to-head matchups, weeks 4–15 regular season, top four make the playoffs in weeks 16–17.

One known difference: Jev receives a short description with each label (for example what counts as a "start"), while the two GLiNER managers receive the label names only, because the Fastino API doesn't yet accept label descriptions. glifan learned what the labels mean from its training data; the untrained Decide control did not.

What the AI managers see

One line of text per player, built from nflverse data. Names are never shown, so the models judge situations, not reputations:

position: WR | week: 3 | home: yes | last3_ppr: 17.4 | trend: up | season_avg_ppr: 25.3 | targets_l3: 6.0 | target_share_l3: 20% | snap_pct_l3: 67% | matchup_rank: 8 of 32 (easy) | implied_team_total: 25.0 | spread: +3.5 | game_total: 53.5 | injury: healthy | practice: full | weather: outdoors

A manager's lineup is its highest-ranked players by P(start) + ½·P(flex) − ½·P(bench), filled into QB, RB, RB, WR, WR, TE and FLEX. Players ruled out on the injury report are never started.

Before the season: a 2025 replay

We replayed the 2025 season, which glifan never trained on, as a 12-team league with the same draft for everyone, putting one manager in each draft seat against eleven season-average managers:

ManagerAvg wins (of 14)PlayoffsTitlesShare of perfect points
glifan (fine-tuned)7.77 / 123 / 1286.9%
Hot hand7.26 / 122 / 1288.1%
Season average–––87.8%
Decide (untrained)5.53 / 121 / 1279.7%

Fine-tuning lifted Decide from 41.5% to 59.3% of weekly start/flex/bench tiers right. glifan won the most head-to-head, though it scored slightly fewer total points than the simple rules, so this season is a real contest.

How glifan was trained

glifan is GLiNER2.5-Decide fine-tuned with LoRA on 14,481 player-weeks from the 2021–2024 regular seasons, labeled by how each player actually finished. It took one dataset upload and one training call on the Fastino API.

If your product makes the same kind of decision thousands of times a day (routing, triage, approvals, moderation), you can fine-tune the same model on your own examples at agent.fastino.ai.