{"id":3216,"date":"2026-10-06T19:05:51","date_gmt":"2026-10-06T19:05:51","guid":{"rendered":"https:\/\/dmarketertayeeb.com\/blog\/tiny-aya-l2-thinker-reasoning-rate-access\/"},"modified":"2026-10-06T19:14:53","modified_gmt":"2026-10-06T19:14:53","slug":"tiny-aya-l2-thinker-reasoning-rate-access","status":"publish","type":"post","link":"https:\/\/dmarketertayeeb.com\/blog\/tiny-aya-l2-thinker-reasoning-rate-access\/","title":{"rendered":"Tiny Aya L2-Thinker: What the 93% Figure Measures"},"content":{"rendered":"\n<p>Cohere&#8217;s <a href=\"https:\/\/cohere.com\/blog\/building-multilingual-bridges\">October 6 announcement<\/a> says Tiny Aya L2-Thinker reasons in the prompt language more than 93% of the time across 60 languages. The <a href=\"https:\/\/arxiv.org\/html\/2609.10445v1\">linked arXiv paper v1<\/a> was submitted September 9, 2026.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Two scores, different questions<\/h2>\n\n\n\n<p>L2 rate means traces are mostly in the prompt language, not accurate answers. MGSM reports accuracy; PolyMath reports a weighted score. L2 rate uses <a href=\"https:\/\/arxiv.org\/html\/2609.10445v1\">FastText<\/a>, falling back to GlotLID for unsupported languages.<\/p>\n\n\n\n<figure class=\"wp-block-table\" style=\"overflow:visible;\"><div aria-label=\"Tiny Aya benchmark results: task score and language rate\" role=\"region\" style=\"overflow-x:auto;outline-offset:3px;\" tabindex=\"0\"><table style=\"width:100%;min-width:660px;table-layout:fixed;\">\n<thead>\n<tr>\n<th scope=\"col\" style=\"white-space:normal;overflow-wrap:anywhere;\">Benchmark (non-English languages)<\/th>\n<th scope=\"col\" style=\"white-space:normal;overflow-wrap:anywhere;\">Task score, mean \u00b1 SD (%)<\/th>\n<th scope=\"col\" style=\"white-space:normal;overflow-wrap:anywhere;\">L2 rate, mean \u00b1 SD (%)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<th scope=\"row\" style=\"white-space:normal;overflow-wrap:anywhere;\">MGSM (34)<\/th>\n<td style=\"white-space:normal;overflow-wrap:anywhere;\">68.0 \u00b1 14.1<\/td>\n<td style=\"white-space:normal;overflow-wrap:anywhere;\">96.5 \u00b1 9.6<\/td>\n<\/tr>\n<tr>\n<th scope=\"row\" style=\"white-space:normal;overflow-wrap:anywhere;\">PolyMath (17)<\/th>\n<td style=\"white-space:normal;overflow-wrap:anywhere;\">11.1 \u00b1 2.5<\/td>\n<td style=\"white-space:normal;overflow-wrap:anywhere;\">94.9 \u00b1 7.7<\/td>\n<\/tr>\n<\/tbody>\n<\/table><\/div><\/figure>\n\n\n\n<figure class=\"wp-block-image\"><img alt=\"Two panels: mean trace-language rates and task scores for MGSM and PolyMath.\" decoding=\"async\" height=\"850\" loading=\"lazy\" src=\"https:\/\/dmarketertayeeb.com\/blog\/wp-content\/uploads\/2026\/10\/tiny-aya-reported-metrics.png\" width=\"760\"\/><figcaption style=\"color:#c4cedc;font-size:16px;line-height:1.6;\">Cohere-reported means; table SDs span language means. PolyMath weights medium\/high\/top 2:4:8. <a href=\"https:\/\/arxiv.org\/html\/2609.10445v1\">Scoring note<\/a>.<\/figcaption><\/figure>\n\n\n\n<p>Paper v1 names Tiny Aya L2-Thinker (3.35B); conditions: one completion per example, 32K context. Unreported: checkpoint hash, temperature, seed and inference hardware. Table 1 lacks sample-level denominators for its non-English means. The figures were not independently reproduced; no weights or dataset bytes were downloaded and no inference was run. <a href=\"https:\/\/arxiv.org\/html\/2609.10445v1\">Paper methods<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Proposed evaluation example (not run)<\/h2>\n\n\n\n<p>Prompt text: \u201c\u00bfCu\u00e1nto es 17 \u00d7 19?\u201d; expected answer: \u201c323.\u201d For an actual response, record the trace-language classifier label, a separate human code-switch flag and the final-answer score. Flag a mixed trace for review even if the answer is correct; a trace that stays in Spanish still fails the task criterion if its answer is wrong. The paper gives no numeric cutoff for \u201cpredominantly\u201d; preregister and report yours. No model response is supplied.<\/p>\n\n\n\n<p><strong>Log for each item:<\/strong> model\/checkpoint hash; dataset commit, subset and split; prompt language; benchmark and scoring version; trace-classifier label; human code-switch flag; final answer and score; actual decoding settings. Keep automatic classification separate from human review, and never copy demo defaults into a paper reproduction.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Access and reproduction (checked October 6, 2026)<\/h2>\n\n\n\n<p>The <a href=\"https:\/\/huggingface.co\/CohereLabs\/tiny-aya-l2-thinker\">model card<\/a> describes a text-only checkpoint trained for reasoning in 44 non-English languages plus English. Model files require Hugging Face sign-in and acceptance of CC-BY-NC-4.0 conditions; Cohere Labs&#8217; <a href=\"https:\/\/docs.cohere.com\/docs\/cohere-labs-acceptable-use-policy\">Acceptable Use Policy<\/a> also applies. The card directs commercial users to Cohere sales. Its authored local-use instructions show a Transformers path.<\/p>\n\n\n\n<p>The <a href=\"https:\/\/huggingface.co\/datasets\/CohereLabs\/tiny-aya-l2-thinker-multilingual-reasoning\">dataset card<\/a> provides a public viewer for 44 non-English subsets and lists train\/test splits, 1.24 GB and 286,376 rows. It names prompts from AM-DeepSeek-R1-0528-Distilled, traces and outputs from gpt-oss-120b, and translations using Command A Translate and DeepSeek-V3. The card marks the dataset CC-BY-NC-4.0.<\/p>\n\n\n\n<p><a href=\"https:\/\/arxiv.org\/html\/2609.10445v1\">Paper Table 3<\/a> lists 286,388; the difference is unexplained. Before reproducing, pin the dataset commit, configuration, subset and split, then record the observed row count.<\/p>\n\n\n\n<p>Cohere&#8217;s <a href=\"https:\/\/docs.cohere.com\/docs\/models\">current model list<\/a> lists Tiny Aya Global, Earth, Fire and Water, but not L2-Thinker. <a href=\"https:\/\/huggingface.co\/CohereLabs\/tiny-aya-l2-thinker\">Hugging Face<\/a> shows no inference-provider deployment for this checkpoint. Cohere&#8217;s article links a <a href=\"https:\/\/huggingface.co\/spaces\/CohereLabs\/tiny-aya-l2-thinker\">Space demo<\/a>, which was not tested. A public L2-Thinker API route is therefore unconfirmed; ask Cohere for the exact model ID if hosted inference is required.<\/p>\n\n\n\n<p>For document extraction rather than multilingual reasoning, see our <a href=\"https:\/\/dmarketertayeeb.com\/blog\/cohere-parse-api-document-intelligence\/\">Cohere Parse API guide<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Metric definitions, scores and access.<\/p>\n","protected":false},"author":1,"featured_media":3214,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[178],"tags":[365,300],"class_list":["post-3216","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","tag-ai-benchmarks","tag-ai-models","has-featured-image"],"_links":{"self":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3216","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/comments?post=3216"}],"version-history":[{"count":1,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3216\/revisions"}],"predecessor-version":[{"id":3217,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/posts\/3216\/revisions\/3217"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media\/3214"}],"wp:attachment":[{"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/media?parent=3216"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/categories?post=3216"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dmarketertayeeb.com\/blog\/wp-json\/wp\/v2\/tags?post=3216"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}