扫码下载
BTC $77,939.81 -0.67%
ETH $2,326.10 -2.76%
BNB $638.22 -0.52%
XRP $1.43 -0.17%
SOL $85.72 -1.67%
TRX $0.3290 -0.10%
DOGE $0.0966 +0.38%
ADA $0.2486 -1.27%
BCH $458.73 -0.88%
LINK $9.30 -0.78%
HYPE $41.11 +0.52%
AAVE $94.14 -0.17%
SUI $0.9417 -1.32%
XLM $0.1755 -1.79%
ZEC $336.65 +5.88%
BTC $77,939.81 -0.67%
ETH $2,326.10 -2.76%
BNB $638.22 -0.52%
XRP $1.43 -0.17%
SOL $85.72 -1.67%
TRX $0.3290 -0.10%
DOGE $0.0966 +0.38%
ADA $0.2486 -1.27%
BCH $458.73 -0.88%
LINK $9.30 -0.78%
HYPE $41.11 +0.52%
AAVE $94.14 -0.17%
SUI $0.9417 -1.32%
XLM $0.1755 -1.79%
ZEC $336.65 +5.88%

DeepSeek 推出 NSA,用于超快速的长上下文训练和推理

2025-02-18 16:37:45
收藏

ChainCatcher 消息,据金十报道,DeepSeek 推出 NSA。

DeepSeek 称,NSA 是一种与硬件一致且本机可训练的稀疏注意力机制,用于超快速的长上下文训练和推理。通过针对现代硬件的优化设计,NSA 加快了推理速度,同时降低了预训练成本,而不会影响性能。

在一般基准测试、长上下文任务和基于指令的推理上,它的表现与完全注意力模型相当甚至更好。

关联标签
关联标签
app_icon
ChainCatcher 与创新者共建Web3世界