The Missing Slice

A Critique of Databricks

The Tyranny of False Intelligence

· Machines & Language · 1,484 words, about 7 minutes

On Memory Management, or The Art of Drowning in an Empty Pool

One might start out on one’s Databricks journey with the naïve belief that a platform whose entire raison d’être is the manipulation of vast datasets might have achieved some elementary competence in memory management. One would be catastrophically mistaken.

Databricks exhibits the memory management capabilities of a goldfish with severe amnesia. Present it with a dataset that should comfortably nestle within your cluster’s RAM allocation and watch with grim fascination as it proceeds to asphyxiate itself with the efficiency of an autocratic regime purging its own functionaries. Your careful tuning of partition sizes? Irrelevant. Your architectural considerations? Dismissed. Your desperate supplications to any deity who might conceivably oversee distributed computing? Unanswered and, one suspects, unheard.

The real obscenity is the platform’s pathological reticence about its own failures. The error messages possess all the illuminating quality of a Delphic oracle translated through three dead languages by a committee of lawyers. Was it the shuffle? Perhaps a broadcast join gone rogue? Some parasitic background process that Databricks spawned with all the discretion of a virus? The platform maintains an omertà worthy of a criminal enterprise. You remain in ignorance, and so, seemingly, does Databricks.

The Caching Debacle, or How to Hoard Garbage While Discarding Treasure

If memory management represents Databricks’ incompetence, then its caching behaviour reveals something more sinister: a kind of aggressive, missionary certainty that it knows your needs better than you do yourself. This is the computational equivalent of a butler who burns your correspondence while carefully preserving every piece of junk mail, insisting all the while that he is acting in your best interests.

The platform will cache, with the fervour of a kleptomaniac at an estate sale, precisely the data you will never access again. Meanwhile, the datasets over which you are actually iterating are treated with the regard one might show to yesterday’s fish. The CACHE TABLE command becomes less a directive than a polite suggestion that Databricks is free to ignore, and very often does.

One attempts to understand what is actually cached at any given moment and finds oneself in a position reminiscent of Kremlinologists during the Cold War, attempting to divine policy from the arrangement of officials in parade photographs. You clear the cache; things remain mysteriously cached. You cache explicitly; your data evaporates like morning dew under a Saharan sun. The whole edifice operates on principles that would make Kafka weep with recognition.

And the reasons for these capricious evictions? One might as well consult the entrails of birds. The documentation offers the usual platitudes about “optimal resource utilization,” which is rather like a pickpocket explaining he was merely redistributing wealth more efficiently.

The Promethean Pretensions: On Being Too Clever by Several Halves

This brings us to the fundamental pathology: Databricks suffers from an acute case of intellectual overreach married to technical incompetence, rather like a failed philosophy student who has discovered computer science and concluded that complexity equals profundity.

The platform drowns in its own abstractions with the enthusiasm of Narcissus admiring his reflection. Every simple task becomes an opportunity for the engineering team to demonstrate their theoretical sophistication — never mind that the theory bears approximately the same relationship to practical utility as medieval angelology does to ornithology. It cannot merely solve your problem, no, it must first construct an elaborate framework to solve a generalised version of the problem you will never have and didn’t know existed.

This is software engineering as performance art, where the comprehension of the audience is not merely unimportant but actually antithetical to the aesthetic.

Adaptive Query Execution: The Platonic Form of Meddlesome Automation

Consider Adaptive Query Execution, the original monument to hubris in this endeavour. In theory it optimizes your queries at runtime. In practice, it is a black box of such profound opacity that one half expects it to demand animal sacrifice before revealing its decisions.

AQE will, with the unsolicited helpfulness of a house guest who rearranges your furniture while you sleep, reorganize your query plan mid-execution. What began as a simple, comprehensible query becomes a Daedalian arrangement of shuffle operations that would make Rube Goldberg weep with admiration and Byzantine bureaucrats nod with recognition. Sometimes this alchemy produces gold. Sometimes it produces the computational equivalent of toxic waste. You will seemingly never be able to predict which.

And when it fails? You are invited to debug query plans that possess all the clarity of Hegelian dialectics translated into assembly language. “Unknown” partitions proliferate like rumours in a totalitarian state. Dynamic operations cascade with the predictability of a drunken mathematician.

Photon: Because Your Epistemological Crisis Needed Deepening

Then we encounter Photon, Databricks’ “native vectorized engine” — a phrase that sounds impressive in the way that “synergistic blockchain solutions” sounds impressive, which is to say, not at all to anyone possessed of critical faculties.

Photon promises to accelerate your queries through the magic of vectorization, and indeed sometimes it does. Sometimes it doesn’t. Sometimes it fails in novel and interesting ways. Sometimes it fails silently, with the discretion of an assassin. Predicting which outcome you’ll receive requires either prophetic abilities or a willingness to engage in the kind of empirical trial-and-error that pre-dates the scientific method.

UDFs work until they don’t. Operations succeed until they fail. Performance improves until it catastrophically degrades. The whole system operates with the consistency of Italian government coalitions. One begins to suspect that Photon’s behaviour is determined by some sort of quantum process, existing in superposition between functional and broken until observed, at which point it collapses into whichever state causes you the most professional embarrassment.

The Emperor’s New Cluster, or The Tyranny of Abstraction

What truly merits ridicule is the mendacious simplicity that Databricks advertises. The platform positions itself as making big data “accessible”, “simple” and “democratized”. What they have in fact constructed is a system so encrusted with layers of abstraction and automation that when the inevitable failures occur, you find yourself as helpless as a medieval peasant watching his crops fail and wondering which saint he failed to properly venerate.

You are not working closer to the metal. You are not gaining understanding. You are, instead, operating through a funhouse mirror that shows you a distorted, optimized, cached, auto-tuned version of reality that bears little or no resemblance to the task at hand.

Traditional Spark, for all its verbosity and requirement for manual tuning, possesses one overwhelming virtue: transparency. You know what it is doing. You know why it is doing it. You can understand the failures and address them. With Databricks, you are a passenger in an autonomous vehicle that refuses to explain its route choices, occasionally takes detours through dangerous neighbourhoods, and has disabled all the manual controls. When you arrive late, crashed, or not at all, the vehicle’s only response is an error code that might as well read “COMPUTER SAYS NO.”

The Verdict: A Meditation on Failed Ambition

What we have in Databricks is a case study in how good intentions, substantial capital, and technical sophistication can combine to produce something genuinely worse than the problem it purported to solve. They have taken Apache Spark (admittedly a challenging framework) and buried it under so many layers of “intelligence” that the result is neither intelligent nor simple, but rather unpredictable and opaque.

The memory issues are not bugs; they are the inevitable consequence of a system that has optimised itself into incomprehensibility. The caching chaos is not a flaw; it is the natural result of a platform that believes it knows better than its users. The whole “too clever” pathology is not an accident it is baked into the product’s DNA.

There exists a profound wisdom in tools that do what you instruct them to do, when you instruct them to do it, without attempting to optimize you into a corner or second-guess your intentions. Databricks has forgotten this wisdom in their Promethean quest to build the clevrest platform in the room.

And so you find yourself, three hours into debugging an out-of-memory error on a cluster that by all rational analysis should have memory to spare, staring at logs that reveal nothing, watching utilization metrics that make no sense, and experiencing what can only be described as computational gaslighting. In such moments, you long not for innovation or disruption or any of the other Silicon Valley shibboleths, but simply for software that does what it claims to do without lying to you about it.

Is this too much to ask? Apparently so. But then, we live in an age where saying that the emperor has no clothes is considered unsophisticated, where complexity is mistaken for depth, and where the ability to generate impressive-sounding marketing copy is valued more highly than the ability to write software that actually works.

Databricks is the platform we deserve, which is to say, not the one we need.

Themes: Machines That Don't Work The Corruption of Language

← All essays