it's 3 am. again. the monitor is the only light in the room and i've been staring at three different ai stacks for the last six hours, and it just hit me: we're not building one future. we're building three. completely separate. completely incompatible.
let me explain.
1. the brain you can own
i've been running gemma 4 on my m4 for the last week. local. no api keys. no rate limits. no "please upgrade to pro." just… mine.
and that's the thing people keep missing. google didn't just release another model. they released the idea that intelligence itself can be something you own. not rent. not subscribe to. own.
i've got it sitting next to ollama, next to my own fine-tunes, doing things that would cost me $200/month on any api. code review at 2 am. summarizing my entire obsidian vault. running inference while i sleep.
the "open weights" movement isn't about altruism. it's about control. when the brain lives on your machine, nobody can take it away from you. nobody can nerf it. nobody can decide that your use case isn't "aligned" enough.
gemma 4 is the brain. and for the first time, the brain is actually yours.
2. the nervous system nobody talks about
apple intelligence is doing something completely different, and honestly? it's more interesting than any chatbot.
i don't want to talk to my ai. i want it to know things. i want it to know that when i say "send that to mom," i mean the photo from tuesday, not the screenshot from slack. i want it to understand that my 3 am coding sessions aren't "anomalous behavior" — they're just… me.
apple isn't building a smarter chatbot. they're building a nervous system. it sits underneath everything — your emails, your calendar, your photos, your habits — and it connects dots that no standalone model ever could.
the foundation models framework just dropped, and i've been playing with it. on-device. private. the model doesn't see your data; your data sees the model. that's a fundamentally different architecture than anything google or openai is doing.
apple intelligence is invisible. that's the point. you don't visit it. it just… lives there. in the background. like a reflex.
3. the engine room
and then there's the part nobody posts about on twitter: the raw experimentation. the lmx (large model experimentation) layer.
while everyone's arguing about which chatbot is best, the real work is happening in labs throwing massive compute at fundamental questions. how does intelligence actually scale? what happens when you push a model past the point where it should stop working? when does a statistical pattern matcher become something… else?
this is where i live most nights. not using ai. breaking it. fine-tuning on weird datasets. running benchmarks that nobody asked for. testing whether a 30b parameter model can outperform a 200b one if you train it right.
lmx is the engine room. it's not pretty. it's not user-facing. but it's where the actual breakthroughs happen — the ones that trickle down into everything else six months later.
the synthesis
here's what's wild: these three layers aren't competing. they're stratifying.
- gemma 4 = intelligence you own
- apple intelligence = intelligence that knows you
- lmx = intelligence that evolves
we're leaving the era of "chat with a bot" and entering the era of "live with intelligence." the ai isn't a destination anymore. it's infrastructure. it's the water in the pipes.
and the people who get this — who understand that these three layers need each other — are the ones building the next decade. not the ones arguing about which model scored 2% higher on some benchmark.
what this means for you
if you're building anything with ai right now, stop asking "which model should i use?" and start asking "which layer am i building on?"
are you building the brain? go open-weights. own it. are you building the nervous system? go deep integration. make it invisible. are you building the engine room? go wild. break things. push limits.
the divergence isn't a problem. it's a signal. the future isn't one ai to rule them all. it's a stack. and the stack is already here.
now if you'll excuse me, it's 4 am and my gemma fine-tune just finished training.
written at 3 am, powered by local inference and too much coffee.
Try it
UNDRDR is my capped archive of 1,000 repos under 1,000 stars — the open-source AI stack included.