Desk

Discover

  • Paper Figure 1: Matthew Effect in RL for LLMs showing pass@1 gains concentrating on easier math, code, and agentic problems (arXiv 2609.13443)

    Mila preprint maps Matthew Effect in RL for LLMs and proposes Never Give Up

    A Mila and Université de Montréal arXiv preprint from Sept. 11, 2026, argues RL for LLMs improves easy problems far more than hard ones (the Matthew Effect) and proposes Never Give Up adaptive sampling; Deepscaler and Manufactoria figures are author claims, and linked code remained 404 as of Sept. 16.

  • Official ZGCM-1 technical-report teaser figure for open 7B math and agentic-search model

    Zhongguancun Academy opens ZGCM-1 for math and agentic search

    On Sept. 11, 2026, Zhongguancun Academy and the Zhongguancun Institute of Artificial Intelligence released ZGCM-1, a fully open 7.39B dense model trained from scratch for math and agentic search with 256K context. Author and Hugging Face card scores, including MATH-500, AIME, HMMT, WebWalkerQA, and a ~4.2x training-efficiency figure, remain company claims; Beckmann and CCTest covered the openness without independent bench replication.

  • Official DeepSeek-V4.1-Flash launch cover art from DeepSeek

    DeepSeek ships V4.1-Flash as an open multimodal model as API docs keep V4-Pro live

    DeepSeek released DeepSeek-V4.1-Flash on Sept. 10, 2026, as API id deepseek-flash with MIT weights, native image input, and a 1 million-token context. After the Sept. 14 cutover window, live pricing and changelog say V4-Pro continues with billing unchanged, while the launch page still shows a superseded route-to-Flash line.

  • Official AuK performance comparison chart across speech generation, editing, enhancement, and separation

    Tencent open-sources AuK, a 1.5B model for instruction-guided speech generation and editing

    Tencent Hunyuan released AuK and distilled AuK-Flash on Sept. 9, 2026, with MIT weights on Hugging Face and ModelScope and a technical report on arXiv (2609.08936). The 1.5B flow backbone unifies zero-shot and instruct TTS, content and acoustic edits, paralinguistic edits, enhancement, and separation behind natural-language instructions. Vendor tables lead several generation and editing benchmarks; independent reproduction is still missing.

  • Official Hugging Face / IBM Research blog card art for Granite Time Series PatchTST-FM-r2

    IBM Granite PatchTST-FM-r2 ships open zero-shot time-series forecasting

    On Sept. 9, 2026, IBM released Granite Time Series PatchTST-FM-r2, a roughly 385M-parameter zero-shot forecaster with an 8,192-step context and 99-quantile output. IBM reports strong GIFT-Eval standing; MustHave.ai says treat the ranking as company-reported pending benchmark merge.