groundy
Developer Tools

Kotlin's Creator Wants Specs, Not Ad Hoc Prompts

CodeSpeak compiles structured English into production code via LLMs. Kotlin creator Andrey Breslav's bet: ad hoc prompting is too ambiguous for serious software development.

Published Updated 10 references
On this page8 sections

Kotlin creator Andrey Breslav built a language where you write what software should do, not how to do it. An LLM handles the rest. CodeSpeak compiles structured English specifications into Python, Go, JavaScript, TypeScript, Kotlin, or Swift. In the project’s own case studies on four open source modules, specs ran 5.9x to 9.9x shorter than the code they replaced, with test pass counts holding steady. Seven months of alpha releases later, the pitch has drifted: the homepage now leads with capturing developer intent from agent chats, and code generation is one part of a larger requirements workflow.

What Is CodeSpeak?

CodeSpeak is a specification language: you describe software behavior in structured English, and an LLM compiles the spec into production code. Andrey Breslav, who led Kotlin’s design at JetBrains through its early releases, is its founder, and the project is his own, not his former employer’s. The first public release was a 0.1.0 alpha on PyPI, and a March 2026 write-up described the project as launching that month (byteiota. “CodeSpeak: Kotlin Creator’s Plain-English Language.” byteiota.com. March 2026); by March 24 the CLI had reached version 0.3.7 (CodeSpeak Blog. “Smarter handling of tests and a separate impl command.” codespeak.dev. March 2026).

The design philosophy sits between two extremes Breslav rejects: traditional programming languages carry too much low-level detail, and ad hoc prompting is too ambiguous. “CodeSpeak is neither a formal language, nor just prompting,” he told The Pragmatic Engineer. It is “designed for engineers, not casual users,” and aims to shrink typical application code by roughly 10x. What remains is “the essence of software engineering”: “only the things the human uniquely knows about what needs to happen, because everything else, the machine knows as well.” (The Pragmatic Engineer. “The programming language after Kotlin.” newsletter.pragmaticengineer.com. 2026)

Terminal window
# Install the CodeSpeak CLI
uv tool install codespeak-cli
# Convert existing code into a specification
codespeak takeover <file>
# Compile specs to production code
codespeak build

How Does CodeSpeak Work?

A spec is a plain-text Markdown file with a .cs.md extension. The CLI sends it to a configured LLM, writes generated code back into the repository, and runs tests. Projects can be mixed: some files spec-generated, some hand-written. The February 2026 release that introduced codespeak takeover shipped mixed-mode options such as --skip-tests and glob-based whitelists controlling which files generation may touch. (CodeSpeak Blog. “First glimpse of codespeak takeover.” codespeak.dev. February 2026)

Here is a real spec, from the project’s Folio demo, a roughly 3,000-line file manager in Go built entirely by prompting:

## Keyboard passthrough
When the terminal panel is focused, only the following keys are handled by
the application itself rather than forwarded to the terminal process:
- Ctrl+T (toggle terminal visibility)
- Ctrl+\ (kill terminal session)
- Ctrl+Up / Ctrl+Down (resize terminal panel)
All other keypresses are forwarded to the shell.

The developer describes behavior and constraints, not implementation steps. That closing rule exists because the author wanted htop, vim, and claude to run unmodified inside the terminal, and the project says the rationale came from a prompt in the author’s Claude Code session rather than from the code. (CodeSpeak Blog. “Intent Recovery: The specs you meant to write.” codespeak.dev. April 2026)

A change is a spec edit. To add a “create folder” action on F7, the demo appends a three-line section to one spec and one row to a keybinding table, then runs codespeak build. CodeSpeak generates the filesystem function, the keyboard handler, the input routing, and the panel refresh. One detail from the vendor’s own build log is worth noting: it reports “No tests found,” so the headline app-scale demo shipped without the test enforcement the case studies rely on. (CodeSpeak Blog. “Modular takeover: from vibe-coded app to spec-driven development.” codespeak.dev. April 2026)

codespeak takeover runs the other direction, converting existing code into a spec. The February tutorial fixed GitHub issue #1468 in Microsoft’s MarkItDown by converting Python to a spec, editing it, and rebuilding: a spec change of +23/-3 lines generated +221/-25 lines of code, about a 10x expansion (CodeSpeak Blog. “First glimpse of codespeak takeover.” codespeak.dev. February 2026). The post is candid about gaps; its own open items include verifying that deleted code can be regenerated equivalently from the spec, and that spec edits produce adequate code changes.

March and April releases extended the workflow. Specs can import each other, and builds process dependencies first (CodeSpeak Blog. “Modular takeover.” codespeak.dev. April 2026). Version 0.3.6 (March 17) taught takeover to read Claude Code session histories, behind an explicit permission prompt recorded per project; the same release made coverage reporting language-agnostic via LCOV, allowed whitelisted Anthropic-compatible providers through ANTHROPIC_BASE_URL, switched the default model to Claude Sonnet 4.6, and added an opt-in cap on per-build spend. (CodeSpeak Blog. “Your intent is everything.” codespeak.dev. March 2026)

The Core Problem: Why Ad Hoc Prompts Break Down

The ambiguity problem is fundamental to natural language, not incidental. The launch thread’s most direct statement of it came from commenter the_duke: “Non-deterministic model output, rapidly evolving LLM versions, and underspecified text create huge amounts of details that code has to make concrete.” (byteiota. byteiota.com. March 2026) A chat prompt leaves those details to the model, and the same request can come back differently across runs, context contents, and model versions.

For exploratory prototyping, that variability is tolerable. For systems that get maintained, upgraded, and audited, it compounds: every model update can silently change what prompt-written code does.

Breslav’s answer is to move those decisions into the spec, where they are visible and versioned, rather than leaving them to the model’s discretion. This differs from prompt engineering, which tries to steer model behavior through phrasing. CodeSpeak tries to shrink the set of questions the model answers on its own.

What the Case Studies Actually Show

CodeSpeak’s case studies converted modules from four open source Python projects. These are vendor-published numbers, recorded here from UBOS Tech’s republication of the vendor’s own case-study page; the LOC method strips blank lines and splits long lines, and the Faker case excludes an ~8,000-line table of Italian municipality codes. (UBOS Tech. “CodeSpeak Introduces Innovative Mixed-Project Workflow Platform.” ubos.tech. 2026)

Module (project)Code LOCSpec LOCReductionTests before → after
WebVTT subtitles (yt-dlp)255386.7x1,241/1,242 → 1,278/1,279
Italian SSN generator (Faker)165217.9x2,216 → 2,229
Encoding detection (BeautifulSoup4)8261415.9x889 → 914
EML converter (MarkItDown)139149.9x165 → 192

Every conversion added tests, 13 to 37 per module, while pass counts held. One nuance the summary glosses: yt-dlp went from 1,241 of 1,242 passing to 1,278 of 1,279, so “tests pass” means the pass rate held, not that every test was green.

The 5.9x to 9.9x range tracks the stated 5-10x target, and the variance is informative. Heuristic-heavy code like BeautifulSoup4’s encoding detection compresses less than a simple converter like MarkItDown’s EML handler, because more edge cases must be spelled out in the spec. The Folio demo landed in the same band: 430 lines of specs for 2,938 lines of Go, about a 7x reduction. (CodeSpeak Blog. “Modular takeover.” codespeak.dev. April 2026)

Why This Approach Is Gaining Traction Now

Abstraction has climbed for decades, and JetBrains’ CEO framed the company’s parallel effort in exactly those terms: “And now it’s time to move even higher.” (InfoWorld. InfoWorld. July 2025) Adoption pressure is real: Stack Overflow’s 2025 survey put weekly use of AI coding tools at 65% of developers, as cited in byteiota’s coverage. (byteiota. byteiota.com. March 2026) And Breslav frames CodeSpeak as a response to growing code complexity in an era when LLM agents increasingly write the implementation. (The Pragmatic Engineer. newsletter.pragmaticengineer.com. 2026)

The Criticisms Worth Taking Seriously

The launch discussion on Hacker News drew 176 points and 146 comments; byteiota documented the thread in detail, and three objections stand out. (byteiota. “CodeSpeak: Kotlin Creator’s Plain-English Language.” byteiota.com. March 2026)

The spec-is-as-hard-as-code problem. The oldest objection in specification work survives the LLM era: a spec precise enough to drive correct generation approaches the complexity of the code it replaces. Commenter lifis sharpened it: specs inevitably omit implementation details like variable names, algorithm choices, and data structures, forcing the model into arbitrary decisions, while over-specifying to prevent that defeats the purpose. Breslav’s counter is that specs need only capture what the human uniquely knows, plus tooling to convert code back into specs so the two stay synchronized.

Non-determinism persists below the spec. “Both formal specifications and generated code would be nondeterministic,” commenter pron wrote. “This doesn’t solve the problem—it just moves it.” Rebuilding from the same spec can produce different code, which makes version-control diffs noisy and, as commenter sensanaty put it, makes proving correctness “basically impossible” when output varies. Spec drift across model updates is the maintenance risk that follows from this.

The bottleneck may not be ambiguity. “The problem with formal prompting languages is they assume the bottleneck is ambiguity in the prompt,” commenter tonipotato wrote. “In my experience, the bottleneck is actually the model’s context understanding, not prompt clarity.” A precise spec does not help if the model lacks the surrounding system context to implement it correctly.

None of these is fatal, and the thread also produced mitigations: seanmcdirmid suggested generating tests from specifications, differential testing, avoiding whole-codebase regeneration on spec changes, and keeping specs in version control alongside generated code. Version 0.3.7’s test enforcement builds in the first of these. The objections do mark a boundary: bounded modules with clear input/output contracts are strong candidates; systems thick with interdependencies and unclear requirements are not.

How CodeSpeak Compares to Alternative Approaches

ApproachAmbiguity handlingDeterminismMaintenance burdenAdoption cost
Natural language promptsNone; the model interprets freelyLow; output varies per runHigh spec-drift riskMinimal
CodeSpeak specsStructured; fewer open choicesMedium; still model-dependentMaintain the spec, not the codeAlpha software; structured writing required
Hand-written codeEvery detail explicitHighStandardLanguage-dependent
Formal methods (TLA+, Alloy)Mathematical specificationHigh; verifiedProven correctnessSteep

CodeSpeak occupies the band between natural language and hand-written code, without the verification guarantees of formal methods. byteiota’s bottom-line advice at launch was to wait: the tool was a 0.1.0 alpha needing production stability and real-world validation before investing time, and the project’s own release posts still carry an Alpha Preview warning. (byteiota. byteiota.com. March 2026) Whether that band is wide enough to anchor a durable workflow remains the open question.

What Practitioners Need to Know

The CLI is BYOK: you supply an Anthropic API key, the default model is Claude Sonnet 4.6, and any whitelisted Anthropic-compatible provider works via ANTHROPIC_BASE_URL; an opt-in cap limits spend per build. (CodeSpeak Blog. “Your intent is everything.” codespeak.dev. March 2026) As of 0.3.7 the command set is build (generation with test enforcement), impl (generation without), test (auto-detects pytest, Jest, Go test, and other runners, and fixes failures), and coverage (adds missing tests). (CodeSpeak Blog. “Smarter handling of tests and a separate impl command.” codespeak.dev. March 2026) Session reading currently supports Claude Code only; other agents are on the roadmap. (CodeSpeak Blog. “Intent Recovery.” codespeak.dev. April 2026)

The credible use case remains bounded, well-understood modules: transformers, validators, converters, API clients. Every published demo fits that shape, and the Folio takeover shows the workflow extending to a whole small application through a browser-based modularization wizard that proposes module boundaries you can refine before any spec is written. (CodeSpeak Blog. “Modular takeover.” codespeak.dev. April 2026) Nothing published tests complex domain logic, state machines, or cross-cutting architecture; keep those manual. April’s intent-recovery work tightened generated specs by auditing each sentence as linked to something you actually said, structural, or unanchored, with unanchored sentences removed: a 126-line spec that transcribed escape-sequence tables collapsed into a 21-line spec of user-facing behavior (CodeSpeak Blog. “Intent Recovery: The specs you meant to write.” codespeak.dev. April 2026).

The maintenance model deserves thought before adoption. You maintain specs, not code, which takes discipline as requirements evolve. Teams that already let documentation rot should treat spec hygiene as a first-class process concern.

And check the project’s current state before committing. As of September 2026, the homepage describes CodeSpeak as an “Agentic Engineering Toolkit” that captures intent from agent chats, turns it into structured requirements mapped to the code, and enforces them on every change; it says CodeSpeak “will keep your intent captured automatically,” and until then offers a web tool, Intent Studio, for recovering intent buried in existing chats. (CodeSpeak. “Software Engineering with AI.” codespeak.dev) The spec-to-code compiler documented in the February-to-April releases has not disappeared from the pitch, but the emphasis has moved from generating code toward keeping requirements mapped and enforced as agents keep editing. For a project whose framing has shifted repeatedly since February, that volatility is itself a finding: evaluate what ships this quarter, not what the alpha posts promised.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. CodeSpeak. "Software Engineering with AI." codespeak.devcodespeak.devAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy