The research caches a teacher model's top token scores and processes KL loss in chunks, avoiding two memory spikes that make long-sequence distillation difficult.
AI News Bank
A continuously updated, source-transparent newsroom for AI news, deals, companies, learning, product changes, rumors, and expectations.
Everything AI, separated by evidence type.
News
Confirmed reporting and analysis across models, policy, safety, infrastructure, and society.
25Deals
Funding, M&A, strategic investments, capacity commitments, and market structure.
17Companies
Platform moves, pricing, partnerships, competitive strategy, and business model shifts.
26Learn
Research artifacts, explainers, resources, benchmarks, and practical learning coverage.
123Changes
Model releases, access changes, APIs, deprecations, product updates, and availability.
7Rumors
Clearly labeled reports and unconfirmed signals with source-quality caveats.
25Expectations
Forecasts, deadlines, upcoming releases, and next checkpoints — never framed as fact.
Newest sourced posts
ALTK-Evolve retrieves a task-specific subset of stored lessons instead of sending an entire agent playbook on every step, reducing tokens in IBM's controlled AppWorld runs.
The 3.1-billion-parameter model adds screen understanding, multi-image input, grounding and function calling, while its speed and benchmark claims remain first-party results.
Earth-observation teams can generate geospatial vectors for a chosen place and time, then take the resulting raster into their own analysis tools.
A 2,150-fact benchmark separates whether a model can reveal a fact in familiar context from whether it can produce that fact when directly questioned.
The open reproduction challenge produced thousands of claim-level logbooks, but its own false alarms show why human review still matters.
A January-to-August Hub analysis separates launch excitement from the smaller, older models that remain embedded in real developer workflows.
The Microsoft Research and Xbox prototype lets persistent characters pursue goals, build memories and coordinate while players influence them through conversation.
The research framework scales below the whole-model level and reports lower GPU and power needs on production traces, but is not a generally available service.
The interview study argues that refusal checks and surface-level output tests can miss how chatbot responses affect vulnerable young people in context.
The August 14 cutover leaves existing runs working but moves new custom-model projects toward a public-preview serverless GPU environment.
The August expansion adds partner actions across meetings, travel, entertainment, music and services, with access split by market, account and Gemini mode.