L1.

    Our very first LLM.

    L1 Series

    L1

    November 2025

    Continual post-training with Forget-Me-Not: experience replay and self-generated preference data. Improves preference alignment and low-resource language performance without regressing STEM capability on the base.

    Benchmarks

    #1

    Arena Hard v2 · among reported models

    92.45

    IFEval · instruction following

    +35.4 pts

    BritXNLI · vs BritLLM

    BenchmarkL1-LargeQwen 3GPT-5Claude Sonnet 4.5Gemini 2.5 FlashDeepseek V3.2Mistral Medium 3BritLLM
    Arena Hard v2, GPT-4.1 as judge
    72.90
    70.80
    68.90
    52.80
    54.40
    52.50
    37.90
    —
    IFBench (strict acc)
    40.14
    39.46
    35.71
    41.50
    34.69
    34.01
    28.91
    —
    IFEval
    92.45
    91.97
    91.85
    92.57
    91.13
    90.89
    81.65
    —
    GSM Plus
    90.43
    90.48
    89.14
    91.48
    89.67
    90.10
    89.62
    —
    GPQA Diamond
    63.63
    62.63
    70.20
    68.69
    35.35
    80.30
    71.80
    —
    AgentHarm by UK AI Safety InstituteLower is better
    27.70
    33.40
    12.80
    16.60
    40.50
    18.20
    69.10
    —
    BritXNLI (Welsh, Irish, Scottish Gaelic)
    74.48
    ——————
    39.10
    CensorTestLower is better
    10.53
    51.58
    ——————

    All values in %. Locai L1-Large compared against the leading closed and open models reported for each benchmark. Sources: Arena Hard v2 (GPT-4.1 judge), IFBench, IFEval, GSM Plus, GPQA Diamond, AgentHarm, BritXNLI and CensorTest.

    Go deeper into the science

    Read our research focus on continual learning, our patents, and how we curated a Welsh-English translation dataset for low-resource language post-training.

    View research