← 返回列表

基座模型训练完成!

楼层 3浏览 95 赞 5发于 2026-10-02 18:57最后回复 2026-10-02 19:04 原帖
U
贺煦修@user80#12026-10-02 18:575 赞 · 71 阅

200M模型现状(当然目前是base model,不是 chat model):

  • checkpoint 已训练到 300000 步(训练目标已完成),不再有 NaN,logits 分布健康。
  • 对 prompt “The capital of France is” 的 top-6 预测是 the / a / Paris / France / in / also,合理。
    生成样本质量(temperature 0.85, top-k 40, top-p 0.95, repetition penalty 1.15):

维度 表现
常识事实 弱但可查——知道 Paris 是法国首都,但编造人口数字、方向不对
英文叙事 有连贯故事情节(Bilin Valley 的 Biz 和 Zara),TinyStories 影子明显
维基百科体 拟态很强——标题、See also / References / External links 结构都能仿写
代码 结构像样但逻辑错误(quicksort 写坏),基本不具备代码能力
论文引用 会把"Attention is All You Need"接续成维基条目风格,内容编造
模型配置

项 值
架构 Decoder-only Transformer,RoPE + RMSNorm + SwiGLU
dim / layers / heads / hidden 768 / 12 / 12 / 3072
vocab 50257(GPT-2 BPE)
seq_len 1024(cos/sin cache 1024×32)
参数量 ~200M(dense)
自定义 kernel V2(fused RMSNorm+residual / RoPE-QK / SwiGLU,HIP .so)

3. 训练配置

项 值
总步数 300,000(w5 起步,后 w4 续跑到完)
batch 8 × 1024 tokens/rank(早期 w5 为 4×1024)
优化器 MuonAdamW:2D 矩阵走 Muon(NS5 正交化动量,fp32→fp16 补丁后),其余 AdamW
lr 1e-3(AdamW),muon_lr 0.02,warmup 64 步,decay=linear,wd 0.1
betas / eps (0.9, 0.95)/ 1e-8,momentum 0.95,NS5 steps=5
同步 k=64 local-SGD 风格,sync_bucket 64MB
数据 formal_sota.bin(GPT-2 BPE 全语料 v5c,~826GB)流式读取
评估 eval_uniform.bin,每 1000 步 fixed 4 batch
断点 每 1000 步分 slot 保存(ckpt_200m_w4.pt.slot{0,1}.rank{0-4}.pt)

4. 训练结果

指标 值
最终步数 300,000
最终 loss(末 50 步均值) 1.7947(ppl ≈ 6.0)
eval_loss(external whole-file) ~1.77-1.84 @ 末尾
实际消耗 token 1.537B/rank × 5 ≈ 7.7B tokens
墙钟 458,021 s ≈ 127 h(w4 段)
吞吐 3,356 tok/s/rank 平均;日志内可见 4,700 → 5,300 → 6,900 tok/s 三段(fp16 NS5 补丁在 ~289k 步生效后提升 ~30%)
显存 ~13.8 GB / 16 GB,rank0 稳定
gradient skip 末段 run 858 次(nonfinite,fp16 动态 loss scale 自动回退,scale 末值 16384)

tok/s 图里 270k 处回落、289k 处跳到 6,900,对应 guardian 重启与 fp16 NS5 上线。

5. 基准评测(step 300000 终版权重)

基准 本模型 参考
WikiText test ppl(json 子集,~50 万 token) 109.7 GPT-2 124M zero-shot ~29(WT103 口径);本模型训练语料分布不同,不可直接对排,但差距主要是分布外泛化能力
生成质量 语法流畅、风格模仿强(维基体/童话体) 事实不可信(数字自相矛盾)、代码逻辑错、无多步推理

图:m200_training.png(loss / ppl / 吞吐 / 显存)

实测结果

PROMPT:

Artificial intelligence is transforming

References

Artificial intelligence is transforming the human brain into an image, and a higher quality of communication. To the best of our knowledge, these changes are not a major factor in the brain's ability to communicate effectively to humans, but can make it easier for people to remember and recognize words they use.

Our team also tried to improve the memory of different subjects with some success, but none of that was done successfully (at

PROMPT:

References

Once upon a time, in a land far away called the "Bilin">Paschia Valley), there lived two best friends named Biz and Zara. They loved exploring their beautiful home, full of tall trees filled with vibrant flowers and ancient trees. One day, they stumbled upon an old map hidden beneath some rocks near the riverbank. It showed them that the region was once considered part of the second most important social and economic

PROMPT:

def quicksort(arr):
 with quicksort.list() as a_t:
 # find the first element of the array, from which to take the last one
 # and reverse it and return it
 return arr[0]
 else:
 # otherwise, just return the last element
 return arr[1]

class QuicksortItem(object):
 def __init__(self,

PROMPT:

According to a study published in Nature, the average annual rainfall for all the different regions was 675 mm.

See also

References

External links

 National Park Service official website
 The National Park Service at Global Environmental Data and Statistics

National Park Service: A History, from the World Bank
 National Park Service, a national nature-based agency
 National Park Service website
 
 

Parks in

PROMPT:

The Transformer architecture, introduced in the paper Attention is All You Need, is a software tool for transforming high-level (or low-level) computer programs into a "a." The name of the tool was first used by the authors of the paper and its author.

References

External links
 
 
A book about Fauche's work, from Czaszczyński's , the official article on the free source
M
mrgg@mrgg#22026-10-02 19:0450 阅

看不懂 :roll_eyes: 给佬点个赞。

S
sydney1@sydney1#32026-10-02 19:0449 阅

强强?!

本页 3 楼,抓取于 2026-10-07 14:00