What Faynt actually is
Faynt is a pair of Transformer-based game-playing policies for Super Smash Bros. Melee: one with 10M parameters and one with 75M 1. Each model controls all 26 characters from a *single* checkpoint, rather than using character-specialist networks 1.
The claimed highlights:
- Universal control: one policy for every character and matchup 1.
- Strong play: after reinforcement learning, the 10M model wins 240 of 244 same-character games (98.4%) against fourteen specialist and multi-character prior releases, with a winning record vs every release 1.
- Competitive latency: on an NVIDIA T4, optimized inference over recorded states averages 5.2 ms per decision for the 10M model and 8.7 ms for the 75M model, excluding the emulator and communication 1.
- Zero-delay performance: against a privately supplied Slippi-AI with zero action delay, the 10M model wins all 68 games tested across two conditioning settings 1.
Because prior public bots keep 21–24 frames of action delay while Faynt uses no added delay, the authors explicitly note they have not isolated the contribution of that latency gap to the win-rate differences 1.
For a systems builder, the core story is:
With careful pretraining, targeted post-training, and evaluation-aware validation, you can get a compact, low-latency, multi-behavior policy to beat a zoo of larger and more specialized agents.
Training pipeline: more than just RL
The paper describes multiple stages that you’d recognize from modern LLM and RLHF pipelines, adapted to a real-time game.
1. Pretraining on ~840k human replays
Faynt is first trained on roughly 840,000 human Melee replays 1. The model’s task in this phase is to predict controller actions, i.e., supervised behavioral cloning from recorded play 1.
The authors explicitly study:
- Architecture
- Optimization
- Scaling
- Hyperparameter transfer
to guide this pretraining step on the replay data 1.
You can think of this as the “base model” phase for a game agent: learn a broad prior over “what humans do” given game states, for all characters, ignoring rewards.
2. Supervised post-training: curricula + distillation
After pretraining, Faynt does not go straight to RL. There’s a supervised post-training stage that combines three ideas 1:
- Rank-based curricula
- Outcome-based curricula
- 75M → 10M distillation
The rank/outcome curricula organize training examples according to performance indicators (like outcomes) and relative rankings; the exact mechanics aren’t detailed in the abstract, but the key point is that the supervised fine-tuning data is *not* sampled uniformly 1.
The 10M model is also trained with supervision from the larger 75M model—classic model distillation where the big model’s behavior helps teach the smaller one 1.
Evidence that this stage actually matters:
- On a 152-game benchmark, the supervised 10M model wins 69.7% of games.
- The pretrained 75M model (before this post-training) wins only 45.4% of games on the same benchmark 1.
So a smaller but better-post-trained model substantially outperforms a larger, merely pretrained one.
There’s a twist: the supervised 10M model has higher held-out controller-prediction loss than the pretrained 75M 1. In other words, it got *worse* at naïvely predicting human actions, but *better* at winning.
To resolve this, the authors use a weighted validation loss for checkpoint selection. That weighted metric lines up with win rates:
“The weighted validation loss used for supervised checkpoint selection agrees with the win-rate ordering of all four pretrained and supervised policies.” 1
For an agent stack, this tells you:
- Pure log-loss on human actions is not a good operational proxy for game performance.
- A task-weighted or behavior-weighted validation metric can recover the correct ranking of models by strength.
3. RL focused on Fox mirror matches
After supervised post-training, Faynt runs a reinforcement learning phase, but it’s restricted to Fox mirror matches 1. The abstract does not specify the exact RL algorithm or reward structure. What we know:
- The core RL training environment is Fox vs Fox.
- This RL stage is applied after replay pretraining, distillation, and curricula-based supervised fine-tuning 1.
Despite the narrow RL environment (one character, symmetric matchup), the resulting 10M model is evaluated as a universal multi-character policy.
This is a concrete pattern you can copy:
- Use wide, general data (840k multi-character replays) for pretraining.
- Use concentrated RL on a specific, high-signal environment (Fox mirrors) to sharpen tactics where they matter most for your evaluation/baseline comparisons.
- Rely on generalization from the pretrained policy to transfer that sharpening to other characters.
4. RL performance and behavioral changes
After RL and supervised post-training, both the 10M and 75M models show consistent behavior shifts 1:
- Take less damage per minute
- Build larger early leads
- Win more often after losing the first life
These are measurable behavioral statistics that correlate with increased competitive strength, not just raw win rate 1. If you’re building automated evaluation tools, these are good templates for “shaped” metrics that show whether your RL is changing the right things.
Evaluation stack and benchmarks
The Faynt authors distinguish between different opponent types and evaluation conditions:
- Public specialist/multi-character releases with added delay
- Fourteen released agents (specialist and multi-character) are used as opponents on their supported rosters.
- These baselines all have 21- or 24-frame action delays preserved from their original setups 1.
- On same-character games versus this pool, the Faynt 10M model wins 240/244 (98.4%), with a winning record against every individual release 1.
The authors emphasize that Faynt uses no added delay, and they have not disentangled the performance effect of that difference 1. If you adopt a similar benchmarking approach, you should:
- Annotate every opponent with its action latency.
- Consider running matched-latency ablations if you want claims strictly about policy quality.
- Zero-delay Slippi-AI baseline
- In a separate evaluation, Faynt 10M is matched against a privately supplied zero-delay Slippi-AI model.
- Across 68 games and two conditioning settings, Faynt wins all games 1.
This addresses the concern that Faynt is only strong because it’s faster-reacting than delayed bots. At least for this one zero-delay opponent, the policy itself appears stronger.
- Internal 152-game supervised benchmark
- Used to study supervised vs pretrained models (e.g., 69.7% vs 45.4% win rate mentioned above) 1.
- Tightly coupled to the weighted validation loss used for checkpoint selection 1.
From an orchestration perspective, this implies:
- A benchmarks-as-a-service layer: you need reproducible match configurations, rosters, and statistics.
- Support for action-delay configuration per agent.
- Recording of fine-grained stats (damage per minute, life leads, comeback rates), not just win/loss.
The Faynt release includes both benchmark suites and a platform for automated model tournaments, along with the model weights 1. That’s a ready-made harness if you’re iterating on alternative policies or evaluation agents in Melee.
Inference performance and deployment
On an NVIDIA T4, the reported latency for Faynt is 1:
- 10M model: 5.2 ms per decision on recorded game states.
- 75M model: 8.7 ms per decision on recorded game states.
These numbers explicitly exclude the cost of running the game emulator and any inter-process communication 1.
For real-time control:
- Melee runs at 60 FPS, so each frame is ~16.7 ms.
- A 5–9 ms budget for policy inference leaves substantial headroom for emulator and glue overhead on a modest GPU (T4).
For an agent platform, that suggests you can:
- Run multi-agent tournaments on a single T4 without saturating compute if you batch or parallelize.
- Optionally spin up ensembles or auxiliary networks (e.g., value estimators) while staying within real-time constraints, as long as you manage asynchronous communication carefully.
The fact that a 10M-parameter policy is both strong and faster than the 75M version gives you a concrete latency/strength trade-off knob. The included 75M→10M distillation pathway 1 is the mechanism that lets you shift along that curve without retraining from scratch.
How this fits into an agentic stack
If you’re building agentic systems, Faynt’s ingredients map almost directly onto a pattern you might already be using for language or tool agents:
- Base model training: behavioral cloning on a large dataset (here, 840k human replays) 1.
- Post-training: supervised fine-tuning with curricula and distillation, optimized for task-level metrics rather than raw likelihood 1.
- Targeted RL: constrained RL on a sub-distribution of tasks that most strongly affect performance comparisons (Fox mirrors) 1.
- Evaluation harness: diverse benchmarks with clear differences in conditions (action delays, conditioning modes, etc.) 1.
You can translate this to other real-time domains (robotics sim, MOBA games, RTS) by:
- Replacing Slippi replays with your game/sim logs.
- Defining “mirror matches” or other focused RL arenas where improvements are easy to measure.
- Maintaining a benchmark registry that includes environment parameters (latency, physics variants) alongside agent versions, as Faynt does with delayed vs zero-delay opponents 1.
What to watch next
From the Faynt abstract alone, several directions are implied:
- Scaling behavior: The authors explicitly study scaling and hyperparameter transfer across 10M and 75M 1. If extended, this could yield practical recipes for when a larger policy is actually better, and how to distill it down efficiently.
- Better checkpoint selection signals: The success of weighted validation loss over plain controller-prediction loss 1 is a hint that we need more task-aware validation metrics for agent models.
- Generalist control policies: A single checkpoint for 26 characters 1 is a small but concrete step towards generalist controllers in richer simulators, where you don’t want to maintain separate policies per tool or body.
Because the weights and tournament platform are open-sourced 1, you can treat Faynt as a reference baseline for:
- Evaluating new RL algorithms in a fixed environment.
- Testing alternative architectures or action representations while holding the evaluation stack constant.
- Probing generalization across characters and matchups.
What is not documented
The abstract and metadata do not establish:
- The exact Transformer architectures (number of layers, heads, context length, tokenization of game state, etc.).
- The specific RL algorithm(s) used, their hyperparameters, or training compute budgets.
- Details of the rank-based and outcome-based curricula (sampling policies, ranking definitions, curriculum schedules).
- How 75M→10M distillation is implemented (loss functions, temperature, online vs offline).
- The precise structure of the automated tournaments platform (APIs, orchestration model, batching strategies).
- The distribution of the 840k replays (skill levels, characters, matchups) or preprocessing pipelines.
- Any human vs agent evaluation or human-in-the-loop experiments.
- Per-character or per-matchup win rates beyond the aggregate figures cited.
All of those would require reading the full paper or associated code/documentation beyond what is summarized in the provided abstract.