Do androids dream of flappy birds?

Do androids dream of flappy birds?
Flappy bird came a long way

Today’s Best Band Ever™ is Mermaid Chunky.

For a change, something kind, sweet, silly and optimistic. I think that's their genre too - something kind, sweet, silly and optimistic, or somethingkindsweetsillyandoptimistic for short.

Difficulty: I'm too young to die


The world

Another week passed, but much quieter this time.

Ed Sheeran tried to pull "let's all be friends and keep music out of politics". First they came, buddy.

Fox News is turning against Trump and joining the boycott of the White House. Hopefully this is the start of something new, but let's not forget that Trump is not the cause but a symptom - populism always rises when times get tough.


A small fable about the first rule of cloud-native

Recently I proposed rules of cloud-nativeness, the first of which declares that a cloud-native workload must

run only parts that are needed at the moment (on demand)

Shortly after the previous modernization was over, a vigilant reader (thank you, Alok!) reported that when they tried to log in, the whole website froze. To be honest, I didn't expect that somebody would want to log in - this blog is very low traffic.

The root cause of the issue was simple. S3 is a slow medium, and a transaction waiting for a write to complete was blocking the entire connection pool of SQLite. One Claude session later, I had a fix that no longer waits for the write to complete. It's a little riskier for a busier setup, but an accepted risk for mine - a small fraction of rare logins might need a retry.

💡
Measured risk, especially when measured in money, is an excellent tool for reliability budgeting.

Imagine a company selling 10 products per year, earning one million on each sale. Each sale requires a successful login. Let's assume the customers are impatient, and a failed login causes the company a lost sale, i.e. a loss of one million.

If the probability of such a failure is 2% (a value one can measure or simulate), then over 10 yearly logins the company can expect 10 * 2% = 20% loss of one sale.

20% of one million is 200,000. This number is the maximum yearly budget that is reasonable to invest in ensuring customers can log in successfully on the first try.

In the case of my blog, on the other hand, a failed login is a missed conversation. A shame, but since this blog doesn't drive any sales, the budget to fix it is zero.

Still, what struck me was that the whole website froze, not only the log-in route.

Which eventually reminded me that my modernization is far from complete, and the first rule is still unmet - the system is still largely monolithic. A monolithic system is easier to deploy and observe, but when one part goes down, it takes all other parts with it. It's really all-or-nothing.

To make matters worse, a monolithic system is a single "thing" from the cost optimization perspective. The database is already stored in S3 (good!), but the main bulk of the application is still running 24/7, whether anyone is visiting the blog.

An obvious optimization direction is to separate the reading part of the blog from the commenting and administrative. The reading part can then become, for instance, a static website served from S3 and updated when I edit an article or somebody leaves a comment. For the writing parts, "scale to zero" can be implemented for the Lightsail containers (AFAIK it doesn't support that out of the box, but nothing is impossible with a little bit of glue code and some Lambda@Edge functions). An even more intriguing option is to rearchitect those parts as Lambda functions.

They say that a system is as strong as its weakest link. It's worse when the whole thing is just one link.


The excitement of Jev

(boomer mode on)

AI is a surprisingly old technology. Here's a rough timeline with distance between items proportional to the number of years:

  1. The term AI was introduced.
  2. Yann LeCun applied backpropagation to a convolutional neural network. Machine learning and neural networks switched from exotic to useful.
  3. ImageNet was created.
  4. AlexNet was introduced, a convolutional neural network that could classify images. Over time, cloud vendors have created standard machine learning offerings for categorization, prediction, sentiment analysis and so on.
  5. The "Attention Is All You Need" paper was published and GPT became possible.
  6. GPT-1 was released. It was amusing and quite bad - we didn't yet know that model size would matter.
  7. "The Bitter Lesson" announced that scale and not handcrafting is all that matters.
  8. First, Stable Diffusion was released in the summer, then ChatGPT, late in autumn.

Since then, we all have seen the explosive adoption of AI. Recently, though, things have started to change with conclusions like making models bigger doesn't make them better anymore.

(boomer mode off)

We are currently in the optimization phase, where things don't get much better, only cheaper, and composite systems like Claude Code or OpenClaw are created. Unbelievably, both were released in 2025, just last year.

That's why when TypeSafe.ai announced that they had created a completely new thing called Jev that is just like an LLM but is not an LLM, the excitement was palpable.

In short, Jev is a decision model that takes the usual prompt, but instead of creating text, reports back which of the decisions listed in the prompt is likely the best.

💡
An LLM is a system that predicts what is the most likely continuation of its input. Its architecture was created for machine translation after it became clear that word-by-word translations are not enough - every language has not only a different vocabulary to explain each phenomenon, but also has a specific way of putting the words together. How words relate to each other, and as a consequence, which words are more likely to follow this or that is what LLMs encode.

A single LLM pass returns a list of next-word candidates and how probable they are. If a chat system always takes the top one, the result is always the same. To make things "spicy" a random one from the top is picked. The amount of randomness is called temperature - the warmer, the more variety.

A GPT chat system takes your input, chooses the next likely word, appends it to your input and repeats the next word search until it's told to stop.

The more I thought about Jev, the more I suspected that the main difference between Jev and a "traditional" LLM is that it simply stops before generating text and reports how likely the provided choices are to be the continuation of the prompt.

An additional point of irritation for me was that unlike previous developments in this area, TypeSafe just released a product without sharing its architecture.

This led me to create Jeb, a proof of concept that takes an unmodified LLM and uses llama.cpp to get the probabilities (logits) instead of generating the text.

To my surprise, the idea worked and my local Jeb Tables was performing reasonably well - it was choosing better options than random chance.

Here, for example, it's playing Doom.

Jeb playing Doom

In the process, I discovered alternative approaches like OpenJev, which inspired me to implement Flappy Bird (the title image) and Doom. Later, I also found that folks at privatemode.ai had implemented an approach similar to mine. That was the final confirmation that my hypothesis stood a chance. As a side effect, we now likely know how the original Jev model works and can run it outside the single offering from TypeSafe.

After all, science and progress have always relied on access to information.

I guess I am an AI scientist now.


Stay tuned, stop aggressors and fight Nazis no matter what they pretend to be or the country they come from.


PS. My code is AI-assisted, but this blog is still stubbornly 100% GPT-free, including 100% hand-crafted, if occasional, em-dashes. I do use AI to identify (but not fix) grammar and stylistic errors.