Gleam doesn't compile to Erlang source anymore
Gleam is a type-safe and scalable language for the Erlang virtual machine and JavaScript runtimes. Today Gleam v1.19.0 has been published, so let's go over what's new.
A new compilation target
Over the last few months Giacomo Cavalieri has entirely rewritten Gleam's Erlang code generator that has an entirely different design, and most notably, outputs a different format. Previously Gleam generated Erlang source code, now it generates Erlang abstract forms.
"Erlang abstract forms" is an intermediate representation used by the Erlang compiler. It is a metadata-annotated tree that represents Erlang syntax, and it is normally produced by running Erlang's tokeniser and parser. It has a binary encoding using Erlang's external term format, and with this binary format we can load our generated code directly, skipping the front-half of the Erlang compiler.
This new Erlang code generator brings several benefits:
- The performance of the compiler has been improved, significantly reducing the build times for Gleam projects running on Erlang.
The location metadata available to the runtime is now accurate to original Gleam source code, rather than to the Erlang source code the compiler would generate. This means, for example, the line numbers in BEAM crash reports and stacktraces are perfectly accurate, while previously they could be inaccurate, only pointing to the nearest function. This metadata could also enable full support for Gleam in debuggers such as edb ,
though we have not done any work on this ourselves.
- The code-quality of the Gleam compiler has been improved. The Erlang code generator was one of the oldest and most stable parts of the Gleam codebase, so while it wasn't causing us any problems it wasn't conforming to the standards and conventions we have today. This new replacement is excellent, and arguably raising the bar for the compiler as whole.
- We never have to hear someone use the word "transpiler" as a pejorative ever again. <sup>1</sup>
Ok, so how fast is it?
I'm going to show you some numbers in a moment, but please remember that benchmarks are always contrived and never tell the full story. This data could be a good introduction or jumping-off point, but a good understanding requires the person to do further research and experience.
The benchmark is based on José
Valim's langcompilebench project. Thank you José! It is a measure of the time
taken to compile 100 modules that each contain 100 functions that return a
"hello world" string. This is practical as this shape of test project can be
easily replicated across different languages to produce the most like-for-like
test projects, but it is limited in what it can tell us as only a small subset
of the features of each language get compiled. In real projects code will be
greatly more varied, and different features will have different compilation
costs in different languages.
The first stage of the code generator rewrite was released in v1.18.0, the previous release, so let's compare v1.17.0 to the newly released v1.19.0. This chart shows the time taken to compile the benchmark project, lower is better.
As you can see, a considerable improvement! This is a full build from scratch, without any caching. Gleam's compilation is incremental, so during typical development it would be much faster as it will not be compiling the entire project.
The original langcompilebench includes only Erlang, Elixir, and Gleam, but I
have extended it an assortment of other popular programming languages, to help
folks get a rough feel for how fast Gleam compiles compared to a language they
are familiar with. I've also included Gleam when compiling to JavaScript.
Here's the results:
Remember: This is a contrived benchmark and is this alone is insufficient to draw any hard conclusions about these languages. That said, these results do suggest that Gleam's compilation is nice and fast, and as a Gleam programmer I can say that Gleam development is very enjoyable, with little time spent waiting for the computer.
Why not target BEAM bytecode directly?
We have moved from from compiling from source code that is fed to the Erlang compiler to an intermediate-representation that is fed into the Erlang compiler, but why not bypass the Erlang compiler altogether? Couldn't we make a BEAM bytecode generator that outperforms the Erlang compiler? Perhaps we could also use Gleam's type information to generate more optimised code too.
While it is possible to achieve these benefits, it's unlikely we would be able to. Unlike Erlang source and Erlang abstract forms, BEAM bytecode is not fixed and unchanging. Each new release of the virtual machine can evolve and improve the bytecode, adding new functionality and sometimes removing functionality that has been made redundant. We would need to commit to forever keeping up-to-date with this evolution, working closely with the Erlang maintainers to be ready for up-coming changes, and to have new versions of Gleam ready for new releases of the virtual machine. It would also be a significant effort to reproduce all the existing optimisations that have been implemented in the Erlang compiler over the decades, even with the help we might have from Gleam's more capable static analysis.
Gleam is a community project supported by sponsorship. We have only a fraction of the finances of languages that are backed by corporations or academic institutions, so we need to think carefully about the most efficient and sustainable ways to use our resources. Gleam is a reliable foundation for software development, every decision we make has to work for years and decades to come. Compiling to Erlang abstract forms is the cost-benefit sweet-spot for Gleam today.
We're also in great company with this decision. Our much-loved older-sibling language Elixir also compiles to Erlang via abstract forms. If it's good enough for Elixir, then it's good enough for Gleam!
Alright, enough about that. There's plenty more in Gleam v1.19.0 to go-over.
JavaScript decision tree assignment optimisation
It's not just the Erlang code generation that has seen some love, there's some good improvements for JavaScript too.
In Gleam flow control is done with pattern matching using a case expression,
and it gets compiled to nested if statements. Because pattern matching is
declarative the compiler is able to reorder and optimise the runtime logic,
using a divide-and-conquer approach to find the right branch as quickly as
possible.
John Downey has improved this process to
generate flatter code, with nested if statements collapsed into a single
condition and fewer intermediate variables. For example, take this Gleam code:
pub fn go(x) {
case x {
Wibble(1, 2) -> 1
_ -> 2
}
}
Previously this small bit of Gleam would compile to this surprisingly large bit of JavaScript<sup>2</sup>:
export function go(x) {
if (isWibble(x)) {
let $ = x[0];
if ($ === 1) {
let $1 = x[1];
if ($1 === 2) {
return 1;
} else {
return 2;
}
} else {
return 2;
}
} else {
return 2;
}
}
But now it generates this:
export function go(x) {
if (isWibble(x) && x[0] === 1 && x[1] === 2) {
return 1;
} else {
return 2;
}
}
A nice improvement, I'm sure you'll agree! Surprisingly this makes little-to-no change to the size of code bundles once minified and compressed (gzip really is magic), but the resulting code has fewer branches for JavaScript engines to optimise.
Thank you John!
List literal optimisation
While they share a syntax in their respective languages, Gleam's immutable persistent list type is not the same as the JavaScript mutable contiguous array type. When Gleam code is compiled to JavaScript any list literal has to be compiled to JavaScript code that constructs the runtime data structures. For example, take this Gleam code:
let numbers = [1, 2, 3]
This would compile to JavaScript code<sup>2</sup> like this, where a JavaScript array is constructed and passed to a function to convert it to a Gleam list.
const numbers = arrayToList([1, 2, 3])
With this release the compile will now generate different code for short list literals, generating more direct code that does not convert from an array.
const numbers = prepend(1, prepend(2, prepend(3, empty)))
With modern JavaScript engines this results in a nice performance improvement, and it is especially impactful for projects that use lots of short lists, like those using the Lustre library. We recorded no improvement for longer lists, so the array-to-list approach is still used for those.
Thank you Giacomo Cavalieri for this!
TypeScript API overloads
When compiling to JavaScript the Gleam compiler will also generate functions for working with the programmer-defined data structures from JavaScript. Alongside this the compiler can also provide TypeScript declaration files, enabling full integration between TypeScript and Gleam in a single project.
One of the functions provided for each custom type is a function to check whether a value is a particular variant or not. For example, given the following type:
pub type Box(value) {
Full(value)
Empty
}
The generated function would have this TypeScript declaration:
export function Box$isFull(value: any): value is Full<unknown>;
The keen-eyed TypeScript programmer readers may notice a problem here. If the
value is already known to be of type Box<number>, then this function can be
used to refine the container type to Full, but the type parameter of number
is generalised to unknown, causing type information loss. This is very
cumbersome.
Giacomo Cavalieri has added an overload to the definition, so the type is preserved whenever possible.
export function Box$isFull<I>(value: Boxlt;I>): value is Full<I>;
export function Box$isFull(value: any): value is Full<unknown>;
Thank you Giacomo!!
Improvements for other build tools
Gleam users will typically use the official build tool that is part of the
gleam executable, but sometimes folks will want to compile and use Gleam code
in other contexts. For example, an Elixir or Erlang programmer may want to use
a dependency package that is written in Gleam. This works fantastically at runtime
as these three BEAM languages have excellent zero-cost interop, but getting to
this point can be tricky, as Elixir and Erlang's main build tools do not have
built-in support for Gleam.
The gleam executable offers several commands that expose compiler
functionality, for use by other build tools. This release includes several
improvements to these commands, with the intent of getting Gleam support in
Elixir's Mix and Erlang's rebar3.
The Erlang virtual machine requires all packages too have a .app resource
file along with the compiler bytecode. Previously these Gleam-supporting build
tools would be expected to provide these, but now gleam will generate the
files for them when compiling to BEAM.
The compile-package command gains a --no-dev flag, which will have the
compiler only load code from the src directory and to skip packages listed as
dev_dependencies.
The export package-information and export package-interface commands can
now print their information to stdout, while previously they would have to
write to a file. Alongside that, the export javascript-prelude and ```
export
typescript-prelude
commands can now write to a file.
Thank you [Rodrigo Álvarez](https://github.com/Papipo) for these additions!
Hopefully we will see Gleam support in Elixir's Mix build tool soon.
## Language server label support
Gleam has an excellent language server built-in, providing IDE functionality to
all editors that support the language server protocol. Possibly the last piece
of major functionality was full support for labels of fields and arguments.
[Alistair Smith](https://github.com/alii) has fixed this, adding support for
go-to-definition, find-references and renaming of labels! Thank you Alistair, I
know many people will be absolutely delighted by this time-saving feature.
## Formatting Gleam code in the browser
There is a WebAssembly build of the compiler, used by [the language
tour](https://tour.gleam.run) and [the
playground](https://playground.gleam.run) to compile Gleam in the web browser.
[John Downey](https://github.com/jtdowney) has added a new `format_source`
function, enabling people to run the Gleam code formatter inside the browser.
We will add this functionality to the playground in the near future. Thank you
John!
## As always, better error messages
Making error messages as clear and as helpful as possible is very important to us. It's all very well for a tool to be nice to use when things are going well, it is when things are going badly that the experience can really help or hurt the programmer's stress levels.
Small accidental syntax errors can be a pain, especially if you're not sure what and where the mistake is.
[0xda157](https://github.com/0xda157) has added a special error for when a git
merge conflict marker is found in the code, and another for when the
`User(..lucy, score: 10)` record update syntax is written with the original
record in the wrong position like `User(score: 10, ..lucy)`. She has also added
an errors for procedural operators that do not exist in Gleam, such as `+=` and
`*=`.
[n0kk23](https://github.com/n0kk23) has added a custom helpful error message
for when `|` is used in pattern matching in a way that is not valid syntax in
Gleam, but is valid in other languages, such as Java.
[Giacomo Cavalieri](https://github.com/giacomocavalieri) has added a helpful
error message for binary operators that are not permitted in constant
expressions, and at the same time he has improved the fault
tolerance[<sup>3</sup>](#fn3) of the compiler in the presence of these mistakes.
[Andrey Kozhev](https://github.com/ankddev) has added extra context to the
error message for when a module tries to use a private type or value from
another module within the same package, letting them know that while it does
exist, it is private. We do not offer this for modules from dependency
packages, to avoid leaking information about code the programmer does not
maintain.
And lastly, [James Dolan](https://github.com/jamesdolan16) has improved the
type checker such that an invalid type alias definition can no longer cause a
cascade of further errors through all usages of the alias.
Thank you all for making Gleam debugging easier and easier.
## And the rest
And thank you to the bug fixers and experience polishers:
[0xda157](https://github.com/0xda157),
[Amr Kadry](https://github.com/Amrkadry),
[Andrey Kozhev](https://github.com/ankddev),
[Giacomo Cavalieri](https://github.com/giacomocavalieri),
[Hari Mohan](https://github.com/seafoamteal),
[Ian Chamberlain](https://github.com/ian-h-chamberlain),
[Jack Programs](https://github.com/jackprogramsjp),
[John Downey](https://github.com/jtdowney),
[Lillian Rose](https://github.com/lillianrubyrose),
[Mar Bloeiman](https://github.com/strawmelonjuice),
[mmustafasenoglu](https://github.com/mmustafasenoglu),
[Naomi Roberts](https://github.com/naomieow),
[Rodrigo Álvarez](https://github.com/Papipo),
[Senthilnathan](https://github.com/ssenthilnathan3),
[Surya Rose](https://github.com/GearsDatapacks), and
[Vivid](https://github.com/absolutely-vivid).
For full details of the many fixes and improvements they've implemented see [the
changelog](https://github.com/gleam-lang/gleam/blob/main/changelog/v1.19.md).
## A call for support
Gleam is not owned by a corporation; instead it is entirely supported by sponsors, most of which contribute between $5 and $20 USD per month, and Gleam is my sole source of income.
We have made great progress towards our goal of being able to appropriately pay
the core team members, but we still have further to go. Please consider
supporting [the project or core team members](/sponsor/).
Thank you to all our sponsors! And special thanks to our top sponsor:
1. "Transpiler" means a compiler that outputs a human-readable format, such as source code. It's a cool sounding word, but most the time people use it to imply that a given compiler is in some way inferior. This is very silly, as there is nothing about compiling to a human-readable format that makes a compiler easier to implement. If you care about that output being nicely formatted it might even be harder than using a binary format. [↩︎](#fnref1)
2. The code has been slightly edited for clarity, but the parts related to this improvement are unchanged. [↩︎](#fnref2)
3. Gleam's compiler is the heart of the Gleam language server, so unlike traditional compilers it needs to be able to provide information about code even when it is in invalid state. If only valid code could be fully analysed then the language server would provide a degraded experience to the programmer when they are half-way-through a refactoring or other large edit. [↩︎](#fnref3)