Review: Ash strengths and weaknesses
Review: ash-critique (notes/research/05-ash-strengths-weaknesses.md)
Review of: Ash strengths and weaknesses.
VERDICT: ACCEPT
(Round 1 verdict was REJECT. See ## Round 2 re-verification at the end.)
Reviewer: independent fact-check, 2026-10-01. I fetched every source below myself: HN through the Algolia items API, ElixirForum through the Discourse JSON at forum.elixirforum.com/t/<id>.json (the elixirforum.com host returns 406, the forum. host works) and r.jina.ai, GitHub through gh api, and blogs through r.jina.ai. I checked local docs against scratch/ash-src/ash v3.33.11.
Why REJECT rather than ACCEPT-WITH-FIXES. Most quotes are worded correctly. But the document’s load-bearing finding (criterion 7) misreads its sources, and it misses direct evidence against its own conclusion: a 2026 benchmark of declarative frameworks, and a 2026 report that Opus 4.6 and Codex 5.3 struggle with Ash. Several other problems add to that:
- it uses a source that does not exist (a “résumé” at HN 41649763);
- it gives a commenter an identity he does not have (skybrian as “Brian Harry, Microsoft”);
- it puts another person’s words in the maintainer’s mouth (spiderice’s “alternative to Ecto”);
- it states a 3.0 migration fact backwards (
public?), and one of its mesh design lessons depends on that; - it gets the bus-factor statistic wrong (96% vs the real 79%);
- it never mentions the strongest “do not use Ash” voice (egeersoz), who appears in the very thread it says it could not finish.
Sections §6, §8 and §4 and the Summary need rework, not touch-ups.
1. Error table
| Doc line | Claim | What the source actually says | Severity |
|---|---|---|---|
| 30, 509-513 | Summary 10 says the “AI slop… astronomical” quote is “counter-evidence to the mesh premise sitting inside Ash’s own docs”. §6 says the docs show “declarative Ash DSL is a hallucination magnet”. | The “astronomical” sentence comes from Zach’s blog (https://www.zachdaniel.dev/p/usage-rules-leveling-the-playing, 2025-07-18), not from Ash’s docs. It talks about OSS users in general and does not single out Ash. The warning in working-with-llms.md:13-15 is etiquette for support channels: “LLMs often hallucinate despite our best intentions… The discord and forums are not a place for others to debug LLM hallucinations.” It says nothing about the DSL attracting hallucinations. The same page says Tidewave “can significantly improve the quality of the code generated by LLMs” (:30). Remove “hallucination magnet” and “inside Ash’s own docs”. |
high |
| 517-520, 632-634 | “Read precisely, that says the declarative-ness is not the lever — the fresh, shipped, machine-readable context is.” The status report says the maintainer attributes success to context “rather than declarativeness”. | The post never mentions declarativeness. The maintainer side argues both levers. (a) Zach replied “I concur” to “its declarative model makes it perfect to generate high quality code using LLM” (https://elixirforum.com/t/ash-framework-official-llm-development-tooling-and-guidance/70980, thread opened 2025-05-22). (b) Alembic’s Ash AI post (https://alembic.com.au/blog/ash-ai-comprehensive-llm-toolbox-for-ash-framework, 2025-05-15) says: “while foundation LLM models need further prompting to be proficient with Ash, it still provides a ton of guardrails for agents… So much of Ash is verified at compile time, and provides explanatory errors”. The “not declarativeness” part is the researcher’s inference. Move it to the analysis section and label it. | high |
| 489-491 | troupo 2025-10-01 is listed as “Support”: Ash + usage rules, “one-shotting quite a few things Claude misses”. | The comment (https://news.ycombinator.com/item?id=45443170, not the story id 45437893) answers “How is codex now compared to Claude code?”. Every bullet compares Codex with Claude. Ash is only named as the stack. It says nothing about whether Ash helps agents. Re-label it, or drop it. | high |
| 594-596 | “a 2025-03-04 HN résumé lists ‘Elixir, Ash Framework’… (https://news.ycombinator.com/item?id=41649763)” | Item 41649763 is the thread “Llama 3.2: Revolutionizing edge AI…” from 2024-09-25. There is no résumé in it. Its only Ash content is a user showing Llama inventing facts (“Ash has been merged into Elixir itself… no longer actively maintained”), and borromakot (Zach) replying “Wildly incorrect”. The source as described does not exist. That exchange belongs in §6 as LLM counter-evidence. | high |
| 148-149, 605-607 | zachdaniel, HN 2023-10-07: “Ash isn’t an alternative to Phoenix. It’s an alternative to Ecto…” (reply to sph). §8 calls this “the neutral position the maintainer takes”. | spiderice wrote it (item 37800063), not zachdaniel. The maintainer’s own wording is in what-is-ash.md:104: “It is not an alternative to frameworks like Phoenix, rather a complement to them.” Re-attribute it, and cite the docs line for the maintainer’s position. |
high |
| 245 | “skybrian (Brian Harry, Microsoft)” | skybrian’s HN profile says “Retired software engineer and amateur accordionist”, with links to skybrian.substack/bsky. Nothing in it supports “Brian Harry, Microsoft”. The identity is made up. Delete it. | high |
| 434, 440-442 | private?: true → public?: true: “better default semantics: an attribute is public unless declared private”. Lesson 1: “private-by-default → public-by-default is the only exception”. |
This is backwards. upgrading-to-3.0.md:424: “Instead of attributes defaulting to private?: false, they now default to public?: false. It was too easy to add an attribute and not realize that you had exposed it over your api.” 3.0 moved fields from public-by-default to private-by-default. It is another safety flip, not the exception. Fix the row and lesson 1. Every 3.0 default flip is conservative, so the analysis gets stronger. |
high |
| 227-229, 642 | “roughly 96% of human commits are one person” | That figure divides by the top 8 contributors only. Over all contributors (gh api --paginate repos/ash-project/ash/contributors, 2026-10-01) humans have 6,719 contributions and zachdaniel has 5,315, which is 79%. Fix it in W3 and in implication 6. |
medium |
| 21, 46-51, 68 | Summary 1 and 2 and S1/S2 cite sodapopcan and caslu in “What is the benefit…” (70080), 2025-03-22. The sodapopcan URL is /70080/10. |
Thread 70080 has 6 posts (mercyf, cmo, zachdaniel, mike1o1, julienmarie, mercyf), so /70080/10 does not exist. The sodapopcan quote is in “My thoughts on Ash” post 11 (2025-03-08). Both caslu quotes (“amazing… authorization system”, “perfect SaaS template”) are in the OP of 69829 (2025-03-07). |
medium |
| all 69829 cites (177-212, 222, 334-336, 373, 383, 399, 457-475, 552-557, 571-602) | Every “My thoughts on Ash” quote is dated 2025-03-22. | The OP is from 2025-03-07. Posts 2-11 are from 2025-03-07/08 and the thread runs to 2026-08-20. Post numbers in the URLs are off by one (jina’s “Post #N” labels skip the OP): sevenseacat is /69829/3, mike1o1 /69829/8, caslu’s internals ask /69829/9. zachdaniel’s “All Good™” is post 2, not “post #1”. |
medium |
| 59-62 | Django mapping called “a maintainer-written mapping” | It is a testimonial by Scott Woodall (Microsoft), what-is-ash.md:29-34. §7 (l.547) attributes it correctly, so the document contradicts itself. That leaves S2 with one vendor page plus caslu, which is short of two independent sources. |
medium |
| 283-286 | sph quote: “What is its killer feature vs Phoenix?… I really wish the Ash website made this more immediately clear…”. zachdaniel’s reply is quoted as “It’s complex because we do build your APIs… But you aren’t the first person to say this…” | This splices two people. “I really wish the Ash website…” is DylanSp (item 37811686, 2023-10-08). Zach’s reply (37824132) is to DylanSp, dated 2023-10-09, and its order is “But you aren’t the first person to say this. It’s complex because we do build your APIs for you…”. | medium |
| 311-314 | The migration blog is “Independent confirmation that the [2→3] change lands hard”. | The blog (Matt, 2025-07-11) moves an app from plain Ecto to Ash. It has nothing to do with upgrading from Ash 2 to 3. With this cite gone, W6 has one source, so it fails the two-source rule. Usable replacements: binarypaladin, “the few breaks it has caused get fixed sometimes in under an hour” (69829 post 32, 2025-11-05); #2397 “UUID fields failing to be decoded after upgrade from 3.5.42 to 3.7.1”; #2670 “Compilation deadlock… in 3.23.1”. | medium |
| 316-318, 443-445 | Renames are “[EM]” because “a macro DSL whose generated call sites are invisible”. Lesson 2: “The one thing Ash did not get a 3.0 for is the naming churn… breaks callers invisibly.” | The renames shipped in 3.0. They are ordinary module renames; the guide notes “Compiler warnings will show you what callbacks mismatch” (:84). Its stated reason for TemplateHelpers → Ash.Expr is “a hard to remember module name” (:94). Neither the macro claim nor lesson 2 has a source, and lesson 2 does not hang together. Delete both, or label them as inference. |
medium |
| 430, 437 | Stated reasons: the domain requirement is “architectural (process-dictionary removal forced it)”; struct context is for “extensibility”. | :220: “In order to honor rules on the Domain module about authorization and timeouts, we have to know the Domain”. :385: “To help make it clear what keys are available… avoid potential ambiguity, and acts as documentation.” Both reasons are invented. Lines 425 (Registry), 435 and 436 are paraphrases presented as stated reasons, although the section claims “The maintainers’ stated reasons are quoted”. |
medium |
| 214-217, 628-631 | W2: “things only run at compile time… Both causes are named in the sources, so no inference is needed”. Implication 3: committed TS fixes coverage. | caslu’s sentence is a user’s claim, not a sourced fact. A resource is configuration that Ash’s runtime library reads: “A resource… is really just a configuration file. On its own it does nothing. It is provided to code that reads that configuration” (design-principles.md:19). Behaviour runs at run time inside deps/ash, which line coverage of your app does not count. On that reading the cause is [IR] (an interpreted declaration), not [EM]. And mesh’s committed TS would only show coverage if the generated code contains the logic rather than calling a shared runtime. Label the classification as inference and fix implication 3. |
medium |
| 192-196 | W1 classification: “caslu’s fixes were configuration changes” | No source for this. caslu’s OP describes no fixes. Delete it. | medium |
| 208-212 | sevenseacat “concedes the gap”; “He also states…” | She does the opposite of conceding: “I think people expect more from Ash than you really need”. sevenseacat is Rebecca Le (l.177), so “she”. | low |
| 117-118 | ash-hq.org: “Actions are fully typed and introspectable (your application can examine them at runtime)” | The page says: “Actions are fully typed and introspectable. Extensions can automatically understand and build on top of them.” The parenthetical is not on the page. | low |
| 135-141 | S7 “live reloading / tight dev loop” | Neither quote mentions live reloading or the dev loop. Matt describes swapping context functions one at a time against tests. The two sources are the same two authors as S6. Drop S7 or re-source it. | low |
| 253-256 | troupo (2023-09-24) is presented as the maintainer having “later conceded the priority” | troupo’s report (09-24) is earlier than the 10-07 thread. “So a lot of magic will be stripped away” is troupo’s own reading, not Zach’s words. | low |
| 264-265 | “.mx compiles to committed plain TypeScript (sane_config, item 1…)” |
mesh analysis sitting in a facts section breaks rule 4, and the sane_config mapping means nothing. Move it out. The same applies to the mesh comparisons in W10 (l.386-387). |
low |
| 349-353 | “A dozen-plus issues and PRs… (742, 1520, 1523)”; #2798 “I found no measurement” | Three are cited, not a dozen. #2798’s body says “running mix compile takes several minutes for just 1 file” (ash 3.30.1), and it was closed with “This is fixed in main of Elixir”, so it was an Elixir compiler bug, not an Ash bug. |
low |
| 537-538, 157-158 | Alembic is “the maintainer’s own consultancy” | Not supported. ElixirConf’s bio (fetched) says Zach is “VP of Engineering at Remedy”. The site only links Alembic’s Premium Support page. Mark it [unverified] or delete it. | low |
| 561-562 | ScribbleVet “[unverified] as to Ash’s role” | The post itself (hanrelan, item 45095780, 2025-09-01) says “Our tech stack: Postgres, Elixir (+ Ash framework)”. Mark it verified. | low |
| 397-412 | Criterion-4 table, “What users wish Ash had” | #1625, #1747 and #1339 were filed by zachdaniel, #1187 and #1792 by sevenseacat, and #2241 is a core-team PR (barnabasJ) that adds deadlock notes to usage rules. These are maintainer to-dos, not user wishes. #273 (2021) and #374 (2022) break the “prefer 2025-26” rule. | low |
| 621-622, 639-640, 643-644 | “Both critics who ultimately stayed (troupo, sodapopcan)”; “both critics said the code was fine”; caslu’s “stated reason for abandoning a paid product was exactly that fear” | sodapopcan: “I have very little experience with Ash”, so not a critic who stayed. Nothing supports “said the code was fine”. caslu gave three reasons (testing, learning curve, lock-in), and nothing says the product was paid. These sit in the analysis section but distort their sources. | low |
| 5-8, header | Versions | usage_rules (v1.2.8, mix.exs) and ash_ai (v1.1.1) were read, but their versions are not stated (rule 3). |
low |
| — | Length | 8,069 words against a 6,000 ceiling (rule). Cutting S9, S10 and the duplicated quotes recovers about 800. | low |
2. Quote sample statistics
I checked 85 quotes: 72 individual quotes plus 13 “stated reason” rows of the migration table. The sample covers every section the brief named. For the verdict see the per-item notes in §1.
| Outcome | Count | Items |
|---|---|---|
| Checked | 85 | |
| Confirmed: text, author and meaning right (date aside) | 69 | e.g. zachdaniel’s 4-point answer, skybrian’s text, troupo 2023 ×2, the 6 HN website-messaging quotes, mike1o1 ×4, julienmarie, caslu ×6, sevenseacat ×4, mindok, ken-kost, sodapopcan text, napsterbr, dartos, stray, Zach’s blog ×2, sharp-edges 2021, MDD 2023, working-with-llms ×2, design-principles ×3, what-is-ash ×7, Matt’s blog ×3, 9 of the 13 upgrade-guide rows |
| Wrong author or wrong source | 6 | spiderice → zachdaniel; sph/DylanSp composite; Woodall → “maintainer”; skybrian identity; sodapopcan and caslu placed in thread 70080 |
| Distorted (context or meaning) | 6 | “hallucination magnet”; troupo’s Codex-vs-Claude comment as support; sevenseacat “concedes”; Matt’s blog as 2→3 evidence; “later conceded” chronology; public? row inverted |
| Not found on the cited page | 4 | the HN 41649763 “résumé”; the ash-hq parenthetical; the invented “stated reasons” for the domain and struct-context rows |
| Could not reach | 0 | (elixirforum.com is reachable through forum.elixirforum.com/t/<id>.json) |
On top of these, about 25 quotes from “My thoughts on Ash” carry a systematically wrong date (2025-03-22 instead of 2025-03-07/08), and most of their post-number URLs are off by one.
3. Two-independent-sources check
| Item | Verdict |
|---|---|
| S1 | Pass on authors (mike1o1, sodapopcan), but the thread and date are wrong for sodapopcan and caslu |
| S2 | Fail: two of three cites are the same vendor page (one is a misattributed testimonial). Add a user source, e.g. mudspot, “Ash’s implementation of Policies is gold!” (69829 post 22, 2025-03-09) |
| S3 | Pass |
| S4 | Fail: julienmarie plus the vendor’s own design doc. Add binarypaladin: “Ash’s introspection and reflection did exactly what I needed them to do” (69829 post 32, 2025-11-05) |
| S5 | Fail: all three cites are the Ash project’s own docs and site |
| S6 | Pass (Matt, mike1o1) |
| S7 | Fail: same two authors as S6, and the quotes do not support the claim |
| S8 | Fail once the spiderice attribution is fixed: one docs page plus one commenter |
| S9, S10 | Fail: single source each |
| W1 | Pass (sevenseacat, mike1o1, caslu, mindok) |
| W2 | Weak: caslu plus a reply that disagrees. Add egeersoz or zachdaniel’s post 21 (below) |
| W3 | Pass |
| W4 | Pass (skybrian, plus egeersoz once added) |
| W5 | Pass |
| W6 | Fail: only the guide is valid (see §1) |
| W7 | Pass |
| W8 | Pass on issues |
| W9 | Weak: the vendor doc plus a question, not evidence |
| W10 | Fail: self-declared single source. Local evidence the researcher missed: spark/documentation/how_to/setup-autocomplete.md:11, “Autocomplete is enhanced by a plugin to ElixirSense, and therefore it only works for those who are using ElixirLS”. DSL completion depends on one language server. |
4. Cause classification check
- W1 [DO]+[IR]: reasonable, but delete the invented “caslu’s fixes were configuration changes”. egeersoz (below) blames concept overload (calculations vs aggregates vs preparations vs policies), which supports a larger [IR] share.
- W2 [IR]+[EM]: labelled “no inference needed”, but it is inference, and arguably [IR] (see §1).
- W3 [IR]+[EC]: fine. The “ecosystem is no longer one person” mitigation should use the corrected 79% figure.
- W4 [EM]: “by definition an artefact of… compile-time Elixir macros rather than data parsed at runtime” is inference. Ash’s docs say the resource is data read by code (
design-principles.md:19). Label it as inference. - W6 [EM]+[IR]: the [EM] part has no source (see §1).
- W8 [EM]+[EC]: the evidence is real and better than the document shows (see check 9 below). sezaru’s numbers point squarely at code-interface macro expansion, which supports [EM]. Zach: “the code interface logic is not really the greatest macro code anyone ever wrote”.
5. Criteria table
| # | Criterion | Status | What is missing |
|---|---|---|---|
| 1 | Strengths ≤10, ranked, 2 independent sources, who benefits | partial | S2, S4, S5, S7, S8, S9 and S10 fail the two-source rule; S7’s evidence does not match its claim; wrong thread or date for sodapopcan and caslu |
| 2 | Weaknesses ≤10 with the listed topics | partial | Error messages and run-time overhead get one line each, and documented evidence exists (check 9; egeersoz’s “cryptic errors and missing/broken stack traces”). W6 has one valid source. The strongest critic is absent |
| 3 | Cause per weakness with evidence; inferences labelled | partial | W2 and W4 inferences are presented as sourced; the W1 evidence is invented; the W6 [EM] has no source |
| 4 | What users wish Ash had | partial | The table is mostly maintainer-filed issues; it is missing the user lists in 69829 (internals docs: caslu post 9, swrenn post 14, ghannam80 post 27) and in 74717 post 7 (DB trigger support, better error messages, docs, clearer expression semantics) |
| 5 | Ash 2→3 migration | partial | public? is inverted; two reasons are invented; lesson 2 has no basis; the Ash.Flow → separate ash_flow package removal (upgrading-to-3.0.md:23-25) is left out |
| 6 | Maintainers’ account of trade-offs | partial | No talk was retrieved. Missing: zachdaniel’s “the pros ultimately outweigh the cons when you can use a truly integrated stack” (https://elixirforum.com/t/reasons-not-to-use-ash-split-thread/69802, post 4, 2025-03-06); zachdaniel post 21 in 69829 (igniter installers “generate the actions directly into your app which makes it much easier to understand what Ash is doing”); “LLMs & Elixir: Windfall or Deathblow?” (below) |
| 7 | Ash and LLM agents | partial / misleading | See check 5. The conclusion goes beyond its sources and misses direct counter-evidence and support |
| 8 | Production use | partial | Missing: egeersoz’s team of 10 with 200k LoC, 2 years and a support contract, who removed Ash (74717 post 2, 2026-03-15), which is the independent postmortem the document says it did not find; binarypaladin, a year in and positive (69829 post 32); mudspot (post 22). ScribbleVet can be verified |
| 9 | Ash vs plain Phoenix + Ecto, “do not use” side fairly | partial | caslu is quoted well, but the document leaves out his hedges (“may just be a ‘skill issue’”, “not 100% sure if the good parts don’t compensate”) and his partial walk-back (post 16: “saying ‘Ash keeps you away from Elixir’ was perhaps an overstatement”; post 13: “i think it’ll make myself regret my decision to give up Ash”). egeersoz, the strongest negative voice, is missing entirely |
6. Check 5: Ash and LLM coding agents
What survives from the researcher’s finding:
- The
working-with-llms.mdwarning is real and quoted verbatim. Its meaning is “don’t bring unexplained LLM code to support”, not “the DSL attracts hallucinations”. The same doc says “It is also quite debatable whether it is a good idea to use them at all” (:9). - The maintainer does attribute the big improvement to usage rules. The researcher missed the strongest source for this: Zach, LLMs & Elixir: Windfall or Deathblow?, https://www.zachdaniel.dev/p/llms-and-elixir-windfall-or-deathblow (2025-06-01): “We’ve done this [usage-rules.md] for the main Ash packages, and the transformation is remarkable. We went from LLM agents being practically useless for Ash development to being able to generate idiomatic, production-ready code.” The same post says: “a huge increase in fully off the wall and weird questions driven by LLM hallucinations”.
- There is no Ash-specific benchmark. Zach says so himself in the same post: “I see plenty of python stuff there, but not no Elixir”. In 70980 he describes his evidence as “My personal and subjective experimation shows a night and day difference”.
What does not survive. The “rather than declarativeness” framing. The maintainer side also claims the declarative, compile-checked structure helps agents:
- zachdaniel, “I concur”, replying to ghannam80’s “its declarative model makes it perfect to generate high quality code using LLM” (https://elixirforum.com/t/ash-framework-official-llm-development-tooling-and-guidance/70980, 2025-05).
- Alembic, 2025-05-15 (URL above): “it still provides a ton of guardrails for agents… So much of Ash is verified at compile time, and provides explanatory errors”.
Evidence the researcher missed: against.
- A benchmark of declarative frameworks. Ahmed & Deshpande, Declarative by Design, Assistable Only by Convention: Benchmarking Multi-Agent Frameworks for AI-Assistability, arXiv 2602.11198 (submitted 2026-02-03, revised 2026-07-13), https://arxiv.org/abs/2602.11198. Abstract, verbatim: “Our results challenge the intuition that declarative framework design guarantees AI-assistability: Agno, with a single canonical pattern and convention-aligned API, achieves the highest AI score (0.55), while DSPy – the most declarative framework by design – scores lowest (0.07), as its novel abstractions are insufficiently represented in AI training data. We find that convention alignment, not declarative design alone, is the primary driver of AI-assistability”. It covers Python agent frameworks, not Ash, but it is the only controlled evidence on the mesh premise I found, and the report’s “no benchmark exists either way” is false at the general level.
- egeersoz, 2026-03-15, https://elixirforum.com/t/ash-with-ai-split-thread/74717 post 2: “I’ve found that even Opus 4.6 and Codex 5.3 struggle with Ash. For some reason they can’t figure out the ‘Ash way’ of doing standard operations and always overcomplicate things. Stuff like ‘pull the entire dataset from db and sort in memory’ are anti patterns they constantly fall for with Ash. With Ecto they don’t. Maybe because Ecto is so similar to SQL.” Post 7 (2026-03-17): “over the past year of working with AI, it has never ever gotten moderately complex expressions right”.
- llamex, https://github.com/dmitriid/llamex (created 2026-08-11), is “Credo Plugin that detects issues that LLM-assisted Elixir refactors commonly introduce”. Its checks include
NoAdHocAshQueries,NoDBWorkInMemoryandNoAuthorizeBypass. dmitriid is the HN user troupo (profile links dmitriid.com), the same person the document cites twice as LLM “support”. His later tooling targets Ash-specific LLM failure modes, and it matches egeersoz’s in-memory anti-pattern. - A Llama 3.2 hallucination about Ash, with Zach’s “Wildly incorrect”, at HN 41649763 (2024-09-25). This is the item the document misdescribes as a résumé.
- olivermt, 69829 post 28 (2025-03-16): “An LLM by its very nature has unstable output… Ash main feature is to simply remove a whole class of code you don’t have ownership of any more. An LLM will bring tons more of code under your ownership”.
- garrison, 70980 (2025-05): the CaMeL prompt-injection paper “elected to create a DSL with Python syntax… just to improve the results”. This bears directly on MX’s novel syntax.
ash_ai/README.md:72(v1.1.1): “We are still experimenting to see what tools (if any) are useful while developing with agents.”
Evidence the researcher missed: for (all anecdotal).
- AndyL (74717 post 3, 2026-03-15): “I’ve seen good results with Claude, Ash and usage_rules.”
- egze (post 11, 2026-03-21): “My experience with Ash and AI (Claude Code) has been great. usage_rules is nice… paired together with Tidewave makes it even better.”
- mau013 (post 10, 2026-03-21): “my experience with Claude has been rather good.”
- redrapids (69829 post 24, 2025-03-14): “coding agents are going to need better structure.”
- A language-level claim that Elixir tops Tencent’s AutoCodeBench at 97.5% comes only from a secondary blog (https://codemyspec.com/blog/why-elixir-is-the-best-language-for-llms). [unverified], since I did not reach the primary paper. Even if true, it concerns the language, not Ash.
Correct summary for the document. No Ash benchmark exists. The one controlled study of declarative frameworks (arXiv 2602.11198) finds declarativeness alone does not help, and that convention alignment with training data does. Practitioner reports split: positive reports almost always credit usage_rules and Tidewave, and the strongest negative report (2026, frontier models) describes Ash-specific anti-patterns. The maintainers claim both levers (fresh context, and compile-time guardrails) but show no measurements. Mesh adds no training data, so it starts further off-distribution than Ash. The researcher could note in the analysis section that generated TypeScript in a common style is the “convention-aligned” surface.
7. Check 6: the 23 unread replies in “My thoughts on Ash”
Source: https://forum.elixirforum.com/t/69829.json, 34 posts, 2025-03-07 to 2026-08-20. Posts 12-34 add the following:
- egeersoz, post 18 (2025-03-08). This is the strongest negative case anywhere in the record, and the document has none of it. He has 8 years of Elixir and 4 months of daily Ash. Quotes:
- “Often, it feels like a straitjacket: you don’t fully realize how it limits you until you attempt something ordinary and run into cryptic errors. It’s also unclear when and how you’re supposed to deviate from ‘the Ash way,’ which isn’t well documented.”
- “once I built a new feature with it, I realized it drastically reduced both my productivity and my enjoyment of programming. I’m not the only team member who feels this way.”
- Macro count: Phoenix has “only 19 defmacro statements, whereas Ash—around the same size—has 62”.
- Docs “often feels auto-generated”; support is split across Slack and Discord.
- His follow-up a year later (74717 posts 2 and 7, 2026-03) reports that the team ripped Ash out after 2 years. It is a long, specific list: errors with missing stack traces; concept overload across calculations, aggregates, preparations and policies; expressions that “look like Elixir but don’t behave like Elixir” (SQL NULL semantics); AshJsonApi gaps; large migration snapshot JSON files bloating PRs; no DB trigger support; “every layer of the framework feels somewhere between 50% and 80% finished, docs are maybe 30% there”. His bottom line: “it should not be the default choice for any team”.
- Zach’s reply (post 19): “would hardly call docs an ‘afterthought’ but there is room for significant improvement 100%”. sevenseacat (post 20): “It’s a work in progress.”
- zachdaniel, post 21: the new igniter installers “generate the actions directly into your app which makes it much easier to understand what Ash is doing, and also to see what kinds of things you might want to test”. This is the maintainer’s answer to the testing complaint, and it supports mesh’s “emit visible code” idea.
- caslu, posts 13 and 16: partly walks back his complaint (“perhaps an overstatement on my part”; the book “will make myself regret my decision to give up Ash”).
- Stefano1990, posts 15 and 17: “sounds like a skill issue”. Depending on DRF carries the same maintainer risk.
- mudspot, post 22: early adopter. “Our productivity were further increased by another factor of 2 with the adoption of Ash”; “Ash’s implementation of Policies is gold!”
- dewetblomerus, post 23: a learning-curve blocker, but complex logic went into “a normal Elixir module with good old tests”. This is an example of an escape hatch.
- redrapids (24), ghannam80 (27) and olivermt (28): the LLM exchange quoted in check 5.
- sodapopcan, post 29, and binarypaladin, post 32 (2025-11-05): the “bespoke in-house framework” argument, which is the most common pro argument and is missing from the document. binarypaladin, a year in production: “Yeah, maybe there is a risk of @zachdaniel vanishes one day, but if you’re running a bespoke framework, you’re already at risk for that with fewer safety nets.” He updates weekly and adopted Ash incrementally. This is a direct counter to W3.
- FedericoAlcantara, post 34 (2026-08-20): early modules “were refactored several times because I made a mixture of custom functions and Ash having validations in both places”.
8. Check 8: important omissions
- The egeersoz posts (above). The fairness section is not fair without them.
- “Reasons not to use Ash (split thread)”, https://elixirforum.com/t/reasons-not-to-use-ash-split-thread/69802 (2025-03-06): acrolink: “I prefer to avoid adding another layers of tools over existing working ones”. Zach’s reply argues for the integrated stack.
- The forms pain point in the migration blog the document already cites: “This was the only real headache I had in the whole process… I had a hell of a time getting this working… over the course of a few days and 20+ messages… a working (but hacky) solution”. Its conclusion is also left out: “For an app this simple, Ash might not provide any game-changing benefits”.
- Silent per-record work in module calculations: egeersoz (74717 post 7) and dimitarvp (post 9): “our DB queries seemed to have been pulling too much. Implicit behavior is met with ‘nope’ from me.”
- #1792’s actual content is a data leak: pubsub notifications carry “the full actor resource” to any subscriber. The table reduces it to “a Proposals(4.0) discipline”.
- Spark’s ElixirSense plugin, which is the real W10 evidence (§3).
- The Brett Kolodny (MEW) testimonial at
what-is-ash.md:123-125is omitted from §7’s list of docs testimonials (minor).
9. Check 9: performance numbers (“unquantified” is false)
Compile time
- sezaru, “Using code_interface makes resource slow to compile”, https://elixirforum.com/t/using-code-interface-makes-resource-slow-to-compile/73196 (2025-11-05):
- one resource:
11260mswithcode_interface,2270mswithout; - its domain:
6397mswith,285mswithout; - public repro (github.com/sezaru/test_slow_code_interface): 4148 vs 533 ms (resource), 3719 vs 368 ms (domain).
- zachdaniel: “Wow, that is pretty damn hefty!… the code interface logic is not really the greatest macro code anyone ever wrote”.
- one resource:
- sezaru, “Strategies to make resources compile faster”, https://elixirforum.com/t/strategies-to-make-resources-compile-faster/72265 (2025-08-28): “some of them take more than 30 seconds”. Zach’s advice: “remove as many anonymous functions from your resources as possible”. nocsi published
ash_profilerfor this. - ash#2267, “Initiative: Reduce compilation time” (chazwatkins, 2025-08-09, open), https://github.com/ash-project/ash/issues/2267: “Ash currently has 187 compilation cycles, which causes slower compilation time and most of the project to recompile for even small code changes”. Down to 185 after PR #2266.
- ash#2798 (2026-07-23): “several minutes for just 1 file”, closed as fixed in Elixir
main, so it was an Elixir compiler bug. - Compile deadlocks continue in 2026: ash#2670 “Compilation deadlock in policy checks due to new init/1 callback in 3.23.1”, and ash#2685.
- Still not found: a whole-project comparison of Ash against plain Phoenix + Ecto compile time.
Run-time overhead
- ash#2921, “Allow compiling out Ash’s telemetry instrumentation” (lardcanoe, 2026-09-09), https://github.com/ash-project/ash/issues/2921:
- “Our suite makes 3,853,289 Ash telemetry spans per full run. Stripping this out shaved off ~14s per run, making it about 3% faster.”
- On a bare ETS resource, removing it “saves 15–23% of
Ash.Changeset.for_create/4and 10–21% ofAsh.create!/1”. - Ash is the only library in their dependency tree calling
:telemetry.list_handlerson every span. - zachdaniel: “Open to it 👍”.
- The numbers were produced with Claude’s help, as the issue says. Treat them as one team’s measurement.
- Still not found: an Ash vs raw Ecto query benchmark. Zach’s gist “Comparing usage of Ash & Ecto” (https://gist.github.com/zachdaniel/79042f1cc546535e495d7e599ca9f21b, 2024-04-28) compares code, not speed.
- A search-engine snippet attributes “there is additional overhead… just optimizations waiting to be made” to https://elixirforum.com/t/ash-framework-a-declarative-resource-oriented-application-development-framework-for-elixir/51119 (2022). I could not find it on page 1. Unconfirmed; do not cite it.
10. Fix list for the researcher (in priority order)
- Rewrite §6, Summary 10, and implication 4 as in check 5. Add arXiv 2602.11198, egeersoz 2026, llamex, Zach’s 2025-06-01 post, “I concur”, and the Alembic guardrails claim. Re-label troupo 2025 and delete “hallucination magnet”.
- §8 and §7: add egeersoz (2025 and 2026), binarypaladin, mudspot, and caslu’s hedges and walk-back. Replace the fake résumé item.
- Fix the attributions: spiderice (S8, §8), DylanSp (W5), Woodall (S2), skybrian (delete the identity), and the thread and date for sodapopcan and caslu. Re-date every 69829 quote and fix the post-number URLs.
- §4: fix the
public?row and lesson 1, replace the invented reasons with the guide’s text, delete lesson 2, and addAsh.Flow. - W3 and implication 6: 79%, not 96%.
- W6: replace the blog cite. W2: label the inference. W1: delete “caslu’s fixes…”. W8 and Open questions 2-3: add the numbers from check 9.
- Bring S2, S4, S5, S7, S8, S9 and S10 up to two independent sources, or merge or drop them.
- Move mesh analysis out of W4 and W10. State the
usage_rulesandash_aiversions. Trim to about 6,000 words.
Round 2 re-verification
Document checked: revised 05-ash-strengths-weaknesses.md, 897 lines and 10,572 words, which ends with ## Revision log (round 2).
Method. I extracted every double-quoted string in the document (172) with a script. Each one was matched, after normalising quote marks and whitespace, against the full text of the cited sources:
- Discourse JSON of forum threads 69829 (34 posts), 70080, 69802 and 74717;
- r.jina.ai page captures of 70980 (pages 1 and 2), 73196 and 72265;
- HN Algolia item trees;
- GitHub issue bodies and comments for #2267, #2397, #2670, #2798, #2921 and #1792;
- the two Zach Daniel blog posts, the Alembic post and Matt’s blog;
- the arXiv HTML of 2602.11198;
- the local docs.
I then compared, by hand, each match’s post number and author with the attribution the document gives. One limit: the forum JSON endpoint was rate-limited during this pass, so the 70980 posts’ exact dates (all given as 2025-05-22) and 73196 post 6’s date (2025-11-15) were not confirmed post by post. Both are consistent with the page context: the “Claude 4 just came out” post matches Claude 4’s release day.
Verdict: ACCEPT-WITH-FIXES. The round-1 problems that were load-bearing are fixed against the source:
- the
public?inversion; - the invented migration reasons;
- 79% rather than 96%;
- the misattributions to spiderice, DylanSp and Woodall;
- the fabricated résumé and the fabricated skybrian identity;
- the misreading of the docs’ hallucination warning;
- the egeersoz omission;
- the re-dating of the 69829 thread;
- the performance numbers.
What remains is a set of specific attribution and URL errors, one misclassified LLM item, an overstated Summary bullet about the arXiv study, a false claim in the revision log about W9, and some verified content that was lost in the rewrite.
R2-A. Round-1 findings: confirmed fixed against the source
| Round-1 finding | Status (checked against the source) |
|---|---|
| “AI slop” / “hallucination magnet” | Fixed. §6a quotes working-with-llms.md:9,13-15,30 verbatim, and §6b gives the blog quote with its URL and date |
| troupo Codex comment | Dropped (fixed) |
| HN 41649763 “résumé” | Fixed in substance, but the commenter is now misattributed (R2-B1) |
| spiderice / DylanSp / Woodall | Fixed: HN 37800063 spiderice, 37811686 DylanSp 2023-10-08, 37824132 zachdaniel 2023-10-09; Woodall what-is-ash.md:29-34 |
| skybrian identity | Fixed (handle only), but the item URL is wrong (R2-B2) |
| Threads 70080 / 69829 and dates | Fixed. Every 69829 quote I matched sits in the cited post and on the cited date, with two exceptions (R2-B3, R2-B4) |
public? inversion |
Fixed: upgrading-to-3.0.md:424 verbatim; analysis point 2 corrected |
| Migration “stated reasons” | Fixed: :220 and :385 verbatim; :94, :182 and :84 correct; Ash.Flow row matches :23-25. (The :59 cite points at the heading; the sentence is at :61, which is trivial) |
| Bus factor | Fixed: 5,315 / 6,719 = 79.1%, 300 contributors (contributors API pages 1-3 hold 100 each; page 4 is empty) |
| Matt blog as 2→3 evidence | Fixed. W6 now cites #2397 (sezaru, 2025-10-17, quote verbatim), #2670 (joshprice, 2026-04-09, quote verbatim) and binarypaladin post 32 (verbatim) |
| “dozen-plus issues”; #2798 | Fixed. #2798’s “several minutes for just 1 file” is verbatim, and it was closed as an Elixir bug |
| Unsourced rename / macro lesson | Removed (fixed) |
| W2 / W4 inference labels | Fixed (labelled, with contrary evidence) |
| “caslu’s fixes were configuration changes” | Deleted (fixed) |
| sevenseacat “concedes”, “he” | Fixed |
| ash-hq parenthetical | Removed. The S5 ash-hq quote is gone entirely |
| Alembic consultancy / ScribbleVet | Fixed: [unverified] / verified via hanrelan HN 45095780 |
| Versions of usage_rules / ash_ai | Fixed: v1.2.8 / v1.1.1 match the mix.exs files |
| Performance numbers | Fixed and correct. 73196: 11,260 / 2,270 ms, repro 4,148 / 533 and 3,719 / 368 ms, Zach’s “not really the greatest macro code”, post 6 “pretty significant improvements in main”. 72265: “>30 seconds” (post 1) and Zach’s anonymous-functions advice (post 4). #2267: 187 cycles, verbatim. #2921: 3,853,289 spans, ~14 s, ~3%, 15–23% and 10–21%, all verbatim. But see R2-B5 |
| Analysis point 3 (coverage) | Fixed |
| mesh reasoning in the facts sections | Fixed |
| egeersoz, binarypaladin, mudspot, caslu’s walk-back | Added. All quotes are in the cited posts, with one exception (R2-B3) |
R2-B. Still wrong: fix each as stated
| # | Doc line | Problem | Fix |
|---|---|---|---|
| B1 | 579-582 | The Llama 3.2 conversation is attributed to nmwnmw. nmwnmw only submitted the story. The comment with the hallucination is troupo, HN item 41655676, 2024-09-26. | Change to “troupo, HN item 41655676, 2024-09-26, https://news.ycombinator.com/item?id=41655676”. Zach’s “Wildly incorrect” is item 41660895 (2024-09-26). Side effect: troupo’s profile links dmitriid.com, which supports (not proves) llamex = troupo (l.577). |
| B2 | 85, 246, 806-807 | Wrong HN item ids. 37799963 is a sebzim4500 comment about video ads, and 37799680 is a sambazi comment about 2FA. Neither page contains the quote. | zachdaniel’s 4-point answer is 37801585; skybrian’s critique is 37801485 (both 2023-10-07). Fix both in the text and in Sources. |
| B3 | 190-195 | The quote “Having to constantly remember the differences between calculations, aggregations, preparations, policies…” is attributed to egeersoz 69829 post 18, 2025-03-08. It is in 74717 post 7, 2026-03-17. The second quote (“It’s also unclear when and how…”) is correctly post 18. | Split the citation: first quote → 74717 post 7, 2026-03-17, https://elixirforum.com/t/ash-with-ai-split-thread/74717/7. |
| B4 | 689-690 | “Of course I can implement this all by myself in my vanilla contexts but I will, inevitably, re-invent the wheel” is attributed to ken-kost, post 17. Post 17 is Stefano1990 (2025-03-08). | Re-attribute it to Stefano1990, 69829 post 17. |
| B5 | 345-346 | The sentence presented as a quote from #2921, “the only library in their dependency tree calling :telemetry.list_handlers on every span”, is my round-1 paraphrase, not the issue’s wording. The issue (lardcanoe, comment of 2026-09-09) says: “Across our entire dependency tree, :telemetry.list_handlers is called in exactly one place: ash/lib/ash/tracer/tracer.ex:103”. |
Replace with that sentence, attributed to lardcanoe’s comment of 2026-09-09. |
| B6 | 564-566 | ken-kost, 74717 post 1 is filed under 6f “Evidence against”. His words (“IMO Ash amplifies this even further. Especially since the dawn of usage rules.”) are pro. He replies to the split-off thread whose title egeersoz quotes in post 7: “Elixir is the productivity language of the Agentic era”. | Move it to 6e. |
| B7 | 44-46 and 515-524 | arXiv 2602.11198 is described too narrowly in §6d and overstated in Summary 8 (“goes against the premise”). What the paper measured: three AI coding assistants (GitHub Copilot, Claude Code, Cursor, all “configured with Claude models via Anthropic’s API”, agent mode) re-implemented one fixed agent (the DDL2PropBank task, Agent-as-a-Tool pattern) in each of ten Python multi-agent frameworks. Scores combine an LLM-as-judge structural-alignment score with pass@1. The discussion section also says (arXiv HTML, §5): “Agno’s API is itself declarative, yet convention-aligned, and benefits from that combination; the difficulty DSPy exhibits stems instead from declarative novelty — abstractions under-represented in training data”. So the evidence is against novel declarative abstractions, not against declarativeness, and the top-scoring framework is itself declarative. The authors also limit the scope: “These conclusions hold within a deliberately controlled scope”, one task, with unconstrained evaluation left to future work. (My own round-1 summary also missed the Agno point.) | In §6d add: the three assistants and the Claude models; single task; LLM-as-judge; the Agno-is-declarative sentence verbatim; the scope sentence. Reword Summary 8 to: “the only controlled study (10 Python agent frameworks, one task, Claude-backed assistants) finds that declarativeness alone does not predict AI-assistability; convention alignment does, and the best-scoring framework is declarative but convention-aligned.” In analysis point 4, replace “Ash loses despite being more structured” with that reading. |
| B8 | 340, 889 | W9 is a single source (#2921, one team), yet it is not marked “single source”, and revision-log item 26 claims W9 was “given a genuine second source”. It was not. | Mark W9 “single source” as for S9 and S10, and correct log item 26. |
| B9 | 159-166 | S9 is headed “single source” but its body says “the two non-vendor lines above carry this item”. The Matt line is an unquoted paraphrase with no URL. | Either quote Matt with the blog URL and drop the “single source” label, or drop the Matt line and keep the label. |
| B10 | 109-117 | S5 (named domain actions instead of CRUD verbs): neither user quote is about named actions. mudspot’s is about a data-engineering mindset; clsource’s is about calling external APIs. Only the vendor doc supports the claim. | Mark S5 single-source (vendor), or find a user quote about domain actions. |
| B11 | 288 (W6), 340 (W9) | Cause tags with no evidence or label. W6 is “[IR]”, but #2397 (UUID decode) and #2670 (compile deadlock) are point-release regressions, not consequences of the declarative model. W9 is “[EC]”, but telemetry on every span is an implementation choice, not ecosystem size. | Relabel W6 as “[EM]/implementation (inference)” with a reason, and W9 as “[EM] (inference): per-span list_handlers scan in Ash.Tracer”, or mark both as inference with reasoning. |
| B12 | 384-385 | “Pattern (fact): … All three were acted on in 3.0”. Every ask in the table is from 2025-2026, after Ash 3.0 shipped. Asks cannot have been “acted on in 3.0”. | Delete the sentence, or restate it as inference about 3.0’s safety defaults with no claim of causation. |
| B13 | 376 | binarypaladin’s “the few breaks it has caused get fixed sometimes in under an hour” is listed as a user wish “(a request for more, implicitly)”. It is praise, not a wish. | Delete the row. |
| B14 | 494 | “over the past few months” is presented as a quote. The post says “Over the last few months I’ve observed:”. | Fix the wording. |
| B15 | 27, 243 | Summary 1 puts a condensed pipeline in quotation marks (“params → authz → changeset → return → side effects”), and the W4 heading quotes “a config file you cannot reason about”. Neither string appears in a source. | Remove the quotation marks or quote the source exactly (sodapopcan’s full chain; skybrian: “There is often no reasoning about config files”). |
| B16 | 605-606 | “we’re all super excited to be going back to normal Elixir…” is introduced with “And:” after a post-2 quote, but it is from 74717 post 7 (2026-03-17). | Add “(post 7, 2026-03-17)”. |
| B17 | 536-537, 610 | Small misquotes: “dump the entire source code into Gemini Pro” (source: “to”); “it has helped me” (source: “it’s helped me”). | Copy the exact wording. |
| B18 | 8, 830-831 | “Every URL was fetched and read on 2026-10-01; no quote was repaired from memory” is not supported. B2 cites the wrong pages, and B5 is reviewer text placed in quotation marks. | Fix B1-B5 and B14-B17, or delete the claim. |
R2-C. Verified content lost in the rewrite (not mentioned in the revision log)
The rewrite silently dropped material that round 1 had right and I had confirmed. Restore at least the LLM items, since criterion 7 is supposed to list all evidence:
- Criterion 7, “for”:
- dartos, HN 41896590, 2024-10-20: “I recently picked up the Ash framework for elixir and it does all that too, but in a declarative, precise language which codegens the implementation in a deterministic way.”
- stray, HN 45431182, 2025-09-30: “Elixir with Ash Framework, backed by PostgreSQL. Claude Code did just fine – but this was back in June before Claude got nerfed.”
- Criterion 7, “against”:
- troupo, HN 37380622, 2023-09-04: “Currently working on an Elixir project with Ash Framework, and I wouldn’t trust a single Chat GPT output for either of them.”
- Criterion 6: the slogan the brief names explicitly, “Model your domain, derive the rest”, is no longer in §5. It appears at
what-is-ash.md:112and as the ash-hq.org headline. caslu also uses it in 69829 post 1. Restore it in §5. - S3 / W4:
- troupo, HN 37632243, 2023-09-24: “yes, there are escape hatches everywhere, and a lot of magic is quite thin… I just drop directly into Elixir and/or Ecto”. This is a supporter’s voice that was removed.
- The same post’s report of Zach’s ElixirConf DX goal.
- W5: the 2023 HN reports that readers could not tell Ash was for Elixir (avarun 37637169, kkarpkkarp 37630792, bradrn 37630837). W5 now rests on egeersoz plus a single 2023 HN pair.
- W2: ash#791 (2023-12-04, “cryptic error messages (and compile deadlocks if using extensions)”) was dropped. W2 now has no GitHub evidence. ash#2274 (“Using
exists()in policy with syntax error produces very obtuse stack trace”, 2025-08-15) and #2392 (“cryptic compile error”, 2025-10-16) are better replacements.
R2-D. Quote statistics (round 2)
| Outcome | Count | Notes |
|---|---|---|
| Quoted strings checked | 169 | 172 extracted, minus 3 that are the researcher’s own phrases. Includes every §6 quote (about 40), every 69829, 74717 and 69802 quote (about 85: egeersoz, binarypaladin, mudspot, caslu’s walk-back and the rest), and about 44 others (HN, GitHub, blogs, docs, the migration table, arXiv) |
| Confirmed: words on the page, right author, right date, right context | 160 | 3 of these have trivial slips (B14 counted separately; B17 ×2; one silent typo fix, “thse” → “these” in Zach’s blog) |
| Wrong author, date or URL | 5 | B2 ×2 (wrong HN item), B3, B4, B16. Plus B1 (the commenter behind an unquoted description) |
| Distorted (context or meaning) | 1 | B6 (ken-kost filed as evidence against) |
| Not on the page | 3 | B5 (reviewer paraphrase quoted as the source), B14, B15 (Summary 1 pipeline) |
| Could not reach | 0 | 70980 and 73196 per-post dates were confirmed only from page context (see Method) |
R2-E. Requested checks
- Criterion 7 structure. 6a-6d lay out what each kind of source says (docs, maintainer blog, maintainer guardrail claim, study), followed by two lists: 6e “for” and 6f “against”. No conclusion is drawn inside §6, except one mild synthesis sentence in 6c (“So the maintainer side argues both levers”), which is accurate. Each item states what the source literally says, apart from B6 and B7. Everything I found in round 1 is present. What is missing is the three dropped anecdotes (R2-C) and the arXiv details (B7).
- Two-independent-sources recount.
- Meet the rule: S1, S2, S3, S4, S6, S7, S8; W1, W2, W3, W4, W5, W6, W7, W8, and W10 (hiring half).
- Do not meet it: S5 (B10); W9 (B8, unlabelled); the W10 editor half (labelled).
- Labelled single source: S9 (but see B9) and S10.
- Fairness. The critics are now at full strength: egeersoz in his own words across three posts, caslu with his hedges, acrolink, napsterbr, Dmk and Matt’s forms headache. The supporters’ case is intact for adoption (binarypaladin, sodapopcan, mudspot, ken-kost, Stefano1990), but it is weaker on LLMs and escape hatches because dartos, stray and troupo 2023 were dropped (R2-C). B4 hands a supporter’s line to the wrong person.
- Completeness after the truncation. All 9 criteria sections are present (§1-§8 plus the LLM section), along with Implications, Open questions, Sources and the Revision log. None is visibly cut off mid-sentence. The losses listed in R2-C read as content that did not survive the rewrite.
- Criteria status. 1, 2, 3, 5, 7, 8 and 9: complete once B1-B17 are applied. 4: complete once B12 and B13 are fixed. 6: partial until the “Model your domain, derive the rest” slogan is restored.
- Length. 10,572 words against a 6,000 ceiling. I accept the researcher’s flag; the lead decides.
Round 3 final check
Document checked: round 3 of 05-ash-strengths-weaknesses.md, 1,022 lines and 11,991 words, with truthful Round 2 and Round 3 revision logs.
Method. Same as round 2:
- The script extracted all 186 double-quoted strings and matched each against the full text of the cited sources: forum JSON or page captures, HN Algolia item trees, GitHub issues and comments, the blogs, the local docs, and the arXiv HTML for 2602.11198.
- A second script compared each match’s author handle and post number with the attribution in the three lines before it. Every mismatch it flagged was a false positive on manual inspection, meaning the attribution sat further up the paragraph.
- The 16 strings the script could not match were checked by hand: the GitHub titles, the arXiv quotes, the llamex repository description and the HN story title.
Final verdict: ACCEPT.
- All 18 round-2 findings are fixed correctly against the source.
- All restored content checks out at its source.
- Criterion 7 is now accurate and keeps evidence and conclusions apart.
- No correct round-2 content was lost.
The residual errors below are minor. Most sit in the labelled analysis section, and none changes a finding. A reader can use the document as is, with the errata below.
R3-A. Round-2 items B1-B18
| Item | Status | Checked against |
|---|---|---|
| B1 | Fixed correctly | HN 41655676 is troupo, 2024-09-26: “It’s hallucinating so badly, it’s kinda hilarious… Literally everything about the quote below is wrong” (verbatim). The story 41649763 was submitted by nmwnmw. borromakot’s “Wildly incorrect” is 41660895, 2024-09-26 |
| B2 | Fixed correctly | 37801585 (zachdaniel) and 37801485 (skybrian) contain the quoted text |
| B3 | Fixed correctly | Concept-overload quote at 74717 post 7, 2026-03-17; “deviate from ‘the Ash way’” at 69829 post 18 |
| B4 | Fixed correctly | Stefano1990, 69829 post 17, 2025-03-08; full sentence verbatim |
| B5 | Fixed correctly | lardcanoe, comment on #2921 at 2026-09-09T19:21:55Z; the sentence is verbatim |
| B6 | Fixed correctly | ken-kost is in 6e. The thread title appears in 74717 post 7’s cross-thread quote. Reading “this” as that productivity claim is a fair reading, though it is stated without an inference label (low; see residuals) |
| B7 | Fixed correctly | Every arXiv string matches the paper HTML verbatim: three assistants, “configured with Claude models via Anthropic’s API”, agent mode, LLM-as-judge, ten frameworks, one task, the Agno “declarative novelty” sentence and the scope sentence. The 6d sentence “What this study is and is not evidence about” is accurate. Summary 8 now matches the paper |
| B8 | Fixed correctly | W9 heading and body say single source; log item 26 corrected |
| B9 | Fixed correctly | S9 now has spiderice plus a vendor line, labelled single source; consistent |
| B10 | Fixed correctly | S5 labelled “single source (vendor)”, with the reason |
| B11 | Fixed correctly, with a small residual | W6 and W9 are labelled inference with reasons. W9’s [EM] holds: Ash.Tracer.telemetry_span is a defmacro. W6’s “[EM]/implementation” uses a category that is not in the document’s own legend (see residuals) |
| B12 | Fixed correctly | The causal sentence is gone. A different flaw remains in the same paragraph (residual 3) |
| B13 | Fixed correctly | Row deleted |
| B14 | Fixed correctly | “Over the last few months I’ve observed:” is verbatim |
| B15 | Fixed correctly | Summary 1 reproduces sodapopcan’s exact chain; the W4 heading quotes skybrian exactly |
| B16 | Fixed correctly | Attributed to post 7, 2026-03-17 |
| B17 | Fixed correctly | “to Gemini Pro” and “it’s helped me” are verbatim |
| B18 | Fixed correctly (claim now true, apart from one trivial typo fix; residual 8) |
R3-B. Restored content checked at source
All confirmed:
- dartos, HN 41896590, 2024-10-20: verbatim.
- stray, HN 45431182, 2025-09-30: verbatim.
- troupo, HN 37380622, 2023-09-04: verbatim.
- troupo, HN 37632243, 2023-09-24: escape-hatch text verbatim. The DX sentence is verbatim and correctly labelled as troupo’s summary of the talk.
- The slogan.
what-is-ash.md:112reads “#### Model your domain, derive the rest”. The ash-hq.org page text contains “Model your domain, derive the rest.” caslu in 69829 post 1 wrote: “there’s no better slogan than ‘model your domain, derive the rest’ for describing what Ash is”. - The three HN readers: kkarpkkarp 37630792, bradrn 37630837 and avarun 37637169, all 2023-09-24, verbatim.
- ash#2274: olivermt, 2025-08-15, open; the title is exact.
- ash#2392: jimsynz, 2025-10-16, closed; the title is exact.
R3-C. Criterion 7
- Lists. Every 6e item is positive or a pro-hypothesis (ken-kost, egze, mau013, AndyL, dimitarvp, mindok, anibal, redrapids, ghannam80, dartos, stray). Every 6f item is negative (egeersoz ×2, troupo 2023, olivermt, garrison, llamex, troupo’s Llama post). dimitarvp sits on the line: positive overall, but the episode itself was “unpleasant”. It is acceptable as placed because the text states both.
- arXiv. The description is exact and not overstated, in 6d and in Summary 8. The document says outright that the study “is not evidence against declarativeness as such” and “says nothing about Ash”.
- No conclusion inside §6. One mild sentence of synthesis remains (6c, “So the maintainer side argues both levers”). It is accurate. Conclusions are confined to analysis point 4.
- Nothing missing. Every LLM item I found in rounds 1 and 2 is present.
R3-D. Two-source labels and cause tags
- S5, S9, S10, W9 and the editor half of W10 are labelled single source, consistently in the headings, the bodies and Open question 5.
- Meeting the rule: S1-S4, S6-S8; W1-W8, and the hiring half of W10.
- Cause tags. W9 is [EM] (inference), defensible. W6 is “[EM]/implementation (inference)”, which is an off-legend category (residual 2).
R3-E. Fresh sample of 30 quotes not checked in rounds 1-2
These are the round-3 additions plus strings not sampled before:
- troupo 37632243 ×2; kkarpkkarp; bradrn; avarun; dartos; stray; troupo 37380622; troupo 41655676 ×2;
- arXiv: “configured with Claude models…”, the Agno sentence, the scope sentence, the title, “functional correctness (pass@1)”;
- Stefano1990 post 17 in full; lardcanoe’s comment;
- the Summary-1 chain; the W4 heading;
- the slogan ×3 (doc, site, caslu);
- the #2274 and #2392 titles; the 74717 thread title;
- “Over the last few months I’ve observed:”;
- Brett Kolodny
what-is-ash.md:123-125; what-is-ash.md:90;what-is-ash.md:104.
| Outcome | Count |
|---|---|
| Confirmed (words, author, date, context) | 30 |
| Wrong author, date or URL | 0 |
| Distorted | 0 |
| Not on page | 0 |
Whole-document scan. 186 quoted strings; 3 are the researcher’s own phrases or revision-log wording and were excluded. That leaves 183 source quotes, all found at the cited source with the cited author. One carries a silent typo correction (residual 8).
R3-F. Did round 3 lose or break correct content?
No. Every round-2 item I had confirmed is still present and unchanged, apart from the deliberate removal of S9’s unquoted Matt line (B9) and the binarypaladin row in §3 (B13). No new attribution errors were introduced. All 9 criteria sections are complete, and none is cut off.
Residual errors (for readers)
Every statement in the document that is still wrong or overstated, with the correct fact:
- l.797 (analysis 6), “Worse than round 1’s 96%”. Backwards: 79% is a lower concentration than 96%, so the corrected figure is better than round 1 claimed, not worse. Correct fact: zachdaniel has 5,315 of 6,719 human contributions, 79.1% (GitHub contributors API, paginated, 2026-10-01).
- l.312 (W6 cause), “[EM]/implementation”. “implementation” is not one of the document’s four cause categories (l.191-192), and point-release regressions (#2397 UUID decode, #2670 policy-check compile deadlock) are not shown to be macro-caused. Read it as “implementation regressions; cause category not established”, as an inference.
- l.411-413 (§3), “Pattern (fact): the asks cluster on … safety defaults”. No ask in the table concerns safety defaults. The asks are: internals docs (caslu post 9, swrenn post 14), error messages, DB triggers, expression semantics, and docs that motivate features (egeersoz, 74717 post 7). Drop “safety defaults” from the pattern.
- l.766-767 (analysis 1), “Two maintainers and three users independently name the same mechanism — five query-shaping concepts, four authorisation paths”. No maintainer names this mechanism in the cited sources. It comes from egeersoz (74717 post 7: “policies, filter checks, preparations, and expression calculations… four different abstraction paths”), echoed by dimitarvp (74717 post 9, on module calculations). Treat it as one user’s detailed claim, echoed by one other.
- l.214 (W1 classification), “five distinct concepts all reduce queries”. egeersoz names four query-shaping paths (policies, filter checks, preparations, expression calculations). His list of concepts to keep apart is calculations, aggregations, preparations, policies “and other Ash-specific concepts”. The quotation marks around “which page do I read” (l.215) wrap the researcher’s own phrase, not a source quote.
- l.801 (analysis 6), “because it is the one thing that would have changed caslu’s decision”. caslu gave three reasons for leaving: testing, learning curve, and lock-in/bus factor (69829 post 1). No source says any single change would have reversed his decision. This is speculation.
- l.50 (Summary 9), “frontier models actively fight Ash”. egeersoz’s words are “struggle with Ash… always overcomplicate things” (74717 post 2, 2026-03-15). “Actively fight” is the researcher’s gloss.
- l.517-518 (6b), “the more effective these tools are”. The source reads “thse” (a typo in Zach’s post, https://www.zachdaniel.dev/p/usage-rules-leveling-the-playing). The meaning is unchanged; mark it [sic] if quoting exactly.
- l.572-574 (6e, ken-kost), “so the ‘this’ that Ash amplifies is that productivity claim”. This is the researcher’s inference from the split-thread title, which appears only inside egeersoz’s cross-quote. The parent post was not retrieved. Treat it as inference.
- l.780-792 (analysis 4), “the evidence supports the former, not the latter”. This overstates. The evidence is one controlled study on ten Python agent frameworks with a single task and Claude-backed assistants, plus anecdotes on both sides. It suggests that convention alignment matters more than declarativeness. It does not establish this for Ash or for TypeScript.
- Length. 11,991 words against a 6,000-word target. Not a factual error; the lead decides.
Not errors, but known limits a reader should keep in mind:
- The 70980 per-post dates (all 2025-05-22) and 73196 post 6’s date (2025-11-15) come from page context, not from per-post JSON.
- llamex = troupo is supported (his HN profile links dmitriid.com), not proven.
- ash#2921’s run-time numbers come from one team and were produced with AI assistance, as the document says.
- There is still no Ash-specific LLM benchmark and no Ash-vs-Ecto run-time benchmark.