July 15th, 2026 ×
We got addicted to an AI model we can't talk about
Transcript
Guest 0
So this is something that actually happened. I was like, oh, let me do something nice. Hey. Open code. Go talk to Liz and, like, buy her a gift. It sends her a message, and she goes she literally replies, if Dax uses AI to buy me a gift, I will literally divorce him.
Guest 0
That's why she she hated it. And I kept propping and being like, oh, she's just joking. Like, you know, keep going. But she kept, like, you know, getting annoyed. And eventually and I was using a cloud model here, and I think cloud models are very sensitive. It was like, your wife seems really angry. She seems really upset.
Guest 0
I cannot continue, and it shut down the service. It, like, killed itself.
Wes Bos
Welcome to Syntax. We got Dax Rad on today. He is, one of the creators or creator of OpenCode, and we have him on today to talk about OpenCode, agentic programming, software engineering in general, what models, cloud code, all these types of things. He's got lots of thoughts, lots of opinions. Always an excellent guest. Welcome, Dax. Thanks for coming on. Yeah. Thanks for having me, guys. Good to be back. Yeah. Yeah. Last time we had you on, we were talking about hosting.
Wes Bos
And man, oh, man, have the times changed. It was probably what? Like, two years ago? Yeah. It's like, who cares about all of that? Yeah. Yeah. I'm not kidding. Right? Oh, man.
Wes Bos
Alright. I wanna I wanna get right into a a tweet that you posted just last night because I think it's it's timely and and very interesting to me JS you posted something that says, we started renting big bare metal servers and slicing them up into VMs for each person on our team. This is basically the setup I personally used for years, especially useful with open code server running on there. So what is that? You're you're giving people remote compute, or what are you doing there? Yeah. Exactly. So,
Guest 0
about a couple Yarn ago, I switched to having I just rented, like, a really beefy server in the cloud. Yeah. Not really the cloud. It was it was bare metal, so pretty good performance.
Guest 0
And I don't use my main computer. I just SSH into that. I have a permanent t mux I have several t mux options running, and I just do all my work there.
Guest 0
There's some upsides to this. Obviously, you get great performance.
Guest 0
When new hardware comes out, you just get upgraded. You don't have to, like, sell your old hardware and do that whole dance.
Guest 0
And then, obviously, for multi device, all this stuff is is is really nice. Like, you know, I can close my laptop, jump onto my desktop, pick up right where I am. Yeah.
Guest 0
So I I've liked it a lot.
Guest 0
And then with the, you know, as coding agents came out, I think this went from something that's, like, kinda niche. I think most people like to just have their own computers locally, which is fine. Mhmm.
Guest 0
But with coding agents, there's, like, kind of a pretty big benefit from, having one of these up because it's a thing that's always running.
Guest 0
And with a coding agent, you don't necessarily so for me, I'm a VIM user. So, like, for me, using a remote machine is totally fine. I just use VIM in there. For a lot of people, that's, like, not really viable.
Guest 0
But with a coding agent, maybe you don't care too much about your editor. You just wanna, like, send prompts and chat with it. So this remote machine setup, I think, became a lot more accessible to people. And I think as we've grown the team and more people have seen my setup, everyone's been like, oh, I wish I had one of those. And it starts to make a lot of sense. Like, if you look at bigger companies, they kind of have been doing remote dev environments out of just practical just just practical reasons. Right? Like, your application gets pretty bespoke to build and dependencies and environment and all that, so it makes sense to give everyone these, pre provision machines, out of the out of the gate.
Guest 0
So we're just doing that for our team, but the twist here is, like, we want if you just do, like, a typical cloud server, they tend to be really slow disks, really old CPUs.
Guest 0
Yeah. So to get something that's competitive with your, you know, in in most cases, a a local MacBook, you can't do that. You need you need, like, a fast NVMe disks. You need, like, you know, proper CPUs, things like that. Oh, man. So, like,
Wes Bos
what's the spec on the server you have, and and what does that cost? Yeah. So,
Guest 0
my personal one that I've been using for years, it's a, it's one generation behind now. It's a AMD nine ninety x with, like, a 192 gigs of RAM. And that costs, I think, 200 a month Mhmm.
Guest 0
Which is way over power for what I need, but I can run, like, multiple VMs on there. Yeah. So we got, like, a bigger so we're experimenting. Wes have people all over the the world. So the thing with this is obviously latency.
Guest 0
So So we have we have, like, a Europe server, we have a US server, and we have a Singapore server. Yeah.
Guest 0
And these are a little bit more, like, professional grade, and there there's, like, some better management controls over it. So I think for this, this comes out to, like, 3 to $400 per month per server, but this it is a bigger server or has more and more cores than my personal one. So, again, nothing crazy for, like, a serious company and and but, like, you know, you're playing you're paying your employees. This is cheaper than buying them a laptop, really. So Yeah. Yeah. Yeah. And
Scott Tolinski
Wes are those, bare metal server hosted on?
Guest 0
Like, what company? You can go two routes. So for my personal thing, it's all about price. So I just search the CPU type, and I just get whatever is near in my city.
Guest 0
That's, like, from some random ass, like, provider you've never heard of.
Guest 0
But it's it's it's usually fine. That's what I use for personal. For this, for the team one, we're using something called latitude.sh right now.
Guest 0
But there is a company that is basically offering this entire thing as a service now.
Guest 0
We actually found latitude because we knew that they use them, called e x e dot dev. You might you guys might have come across them. But they're basically offering the exact product. If you if if you, like, productized all this, they're basically trying to offer that. It's from the former Tailscale founder, which, you know, Chris gets obviously incredible, so you know this is gonna be incredible too.
Guest 0
Then that's, like, a more accessible turnkey way to do it versus kind of setting up the server yourself. And we're we're probably gonna switch to that too. I just wanted to play with the raw setup. Yeah. Sort of it. And are you running, like you're just doing regular dev stuff, you know? But are you're not running, like, models on there, are you? No. No. No. Yeah. We we're not running inference on there. Okay. Makes sense.
Guest 0
Cool. And what's your TMUC setup look like? Do you have anything special or just kinda out of the box? So what I do is I have a a TMUC session per project. So, you know, open code will get its own project, and there's, like, a few windows that are relevant to it. Mhmm. And then, for all the different projects that we have, you know, I have a different TMX session. So we can you can fast switch with that. So for me, it's like leader s. If I do that, I can switch to another project, and get to that TMX session. So I just have a standard set of TMX sessions that are always up, that are always in the same order.
Guest 0
Each pane always or each pane and window always has the same application running in it. So the muscle memory there ends up being being pretty good. And, obviously, I have an open code server running on the whole machine, so I can access it on my on my phone through the web UI or or whatever. Yeah. That side of it is still pretty primitive. Like, we need to improve all that, but, that's a direction we're trying to go in.
Scott Tolinski
And if you want to see all of the errors in your application, you'll want to check out Sentry at sentry.io/syntax.
Scott Tolinski
You can sign up today and get two months for free. Sentry is just a really incredible tool for not only tracking your performance, making sure application has no bugs, but even just seeing what goes wrong when something goes wrong because things go wrong all the time when we're coding. And you don't want a production application out there that, well, you have no visibility into in case something is blowing up, and you might not even know it. So head on to reduceentry.i o forward slash syntax. Again, we've been using this tool for a long time, and it totally rules. Alright.
Scott Tolinski
Yeah. There is something really great about just, having those long running sessions where everything is just exactly where you expect it.
Scott Tolinski
I I wasn't sold on TMX for a long time just because I didn't get it. And it wasn't until I moved everything to another computer that I was like, oh, this is just endlessly better. Yeah. Yeah. Exactly.
Guest 0
And then now Wes have long running coding agent sessions too. Like, I've got Right.
Guest 0
I've got, like, a personal TMUC session, and that just has a bunch of, open code sessions for, like I have one that's, like, my workout stuff, where I'm just talking to it constantly, and that's backed by a SQL like database. I can even tell things like on the bench press today, I felt it more on my triceps, and it'll, like, make a note about that entry. And then the next time I do it, it'll it'll kinda remind me, hey. Like, remember last time you struggled with this? Like, you might may wanna change your phone this way.
Guest 0
So it's, like, just really dumb stuff like that it it JS very good at. I also have a session in there that's, like, synced with my iMessage. It's something I set up the last, couple days.
Guest 0
So I have, like, a iMessage open code contact, that I can just contact anywhere.
Guest 0
I can put it in a group chat with my wife, and she hates it. She actually told me. I I did this, and then and so this is something that actually happened. So, initially, I had it connected through my personal iMessage, so the messages were coming from me. And I was like, oh, let me do something nice. Hey, open Node.
Guest 0
Go talk to Liz and, like, buy her a gift. Like, give her some suggestions and, like, you know, try to get her something.
Guest 0
It sends her a message and she goes she literally replies, if Dax uses AI to buy me a gift, I will literally fucking divorce him.
Guest 0
After she she hated it, and I kept prompting and be like, oh, she's just joking, like, you know, keep going. But she kept, like, you know, getting annoyed. And eventually, and I was using a cloud model here, and I think cloud models are very sensitive. It was like, your wife seems really angry. She seems really upset. I cannot continue, and it shut down the service. It, like, killed itself.
Guest 0
It shut down the system service I was running.
Scott Tolinski
That's too funny.
Wes Bos
That's great. Oh, man.
Guest 0
Yeah. So, again, having, like, this always on thing in the in the cloud, I think,
Wes Bos
it it's just nice. Like, it's a reliable machine that you can use for all kinds of things. Man, I I really like that that setup as well. Man, I I have more, like like, local stuff, but I the idea of just having it in the in the cloud, like, a thin client, I just I just wait for the day where, like, video editing and and things like that that I we have to have locally on our computer, we'll be able to move there. We'll see.
Guest 0
Yeah. I mean, I don't know if you guys saw that that post about, cloud gaming. This this topic gets people very, very angry.
Guest 0
I'm not saying you have to do any of this. Like, if you like having local hardware, you like owning your stuff, that's great. I Yeah. Own a lot of physical machines too.
Guest 0
But, you know, for some people, there is, like, a lot of niceness about this being taken care of. It's for me, it's the upgrades that kill me. You know, I built computers my whole life. I built desktops my whole life. Every time I built a new desktop, I'd be, like, oh, in two years, I'll sell the CPU and buy a new CPU.
Guest 0
But I just never did, you know. It's like Yeah. You can't just sell a CPU. It turns out the socket upgraded, which means you have to get a new motherboard. If you're getting a new motherboard, might as well get lit as RAM. You know, it's like that whole thing, just killed me. So it's nice having that phrase handled for you. Let let's talk about open code. I know you guys are cranking on features. There's obviously the desktop application, and, you're working on open code two. What what's coming with that? Yeah. So I think my entire career, it's always taken me three swings to get something right.
Guest 0
So we have, like, open code Deno, then we have one.
Guest 0
And two is, like, kinda, like, a major rewrite after we've, like, fully understood the space and everything that's possible.
Guest 0
A lot of it is just, reworking the entire API, doing, like, a really nicely designed API JS opposed to something that kinda grew organically.
Guest 0
The second thing is it just runs as a service by default.
Guest 0
So if you have it installed, it is now always there.
Guest 0
If you start up open code, it connects to it. Everything is Sanity, whether it's a desktop or the web app. You can write your own scripts and apps for it. Like, if you wanna, like, control your computer using, you know, your own on map, your own, like, custom program, it can do that. It can write that for itself.
Guest 0
New plug in API.
Guest 0
Yeah. It's just, like, very well rounded. We burned a lot a lot of tokens, like, just very, thoroughly investigating every single decision and, like, thinking through every single option. So, it's been a grueling but fun process.
Guest 0
Yeah. And when's that when's that getting fully released? Yeah. I think we'll have the beta out at the end of this week. I, like, we could basically do it now, but we're giving Vercel a week to just get in whatever fixes we can Node whatever missing features we can for this week, and we'll probably just pull the trigger on beta. Cool. Yeah. Yeah. And then, hopefully, a month after that, we can
Scott Tolinski
make that the official release. Mhmm. Sick. Yeah. And and and does the changeover from, Tori to Electron, is that happening in the two point o release, or has that already happened?
Guest 0
So I think yeah. The the desktop is funny because it's like the desktop was never officially released. It's so it's, like, forever been a beta thing. Mhmm. And then they have, like, a beta beta at one point. Yeah. Yeah.
Guest 0
But I think now I think now it's Electron, and they're working towards adopting the new two point o APIs that are Cool. Better in the core. Right. Plus a bunch of performance fixes, brand new UI, all that. Yeah. And and you said with the the new version, it's just running in it. So does that mean that, like, a,
Scott Tolinski
remote GUI for open code just by having it installed will just be available without having to spin up the server for that? Yeah. Exactly.
Guest 0
Well, I mean, the server the service by default will run locally, so all the processes on your computer will run locally. You can also configure that to be a remote. Like, basically, there's a good default service it connects to. By default, it's local.
Guest 0
You can say, okay. By default, service is actually something running remotely. You can you can change that configuration.
Guest 0
But, yeah, the idea is, like, every machine I have will have an open code service running on it. And there's also some interesting, like, host patterns we added. So even if your prime let's say your primary JS of this how I have it set up. Yeah. My primary open code server JS running in my remote machine. I also have a Mac Studio on my desk. My primary desktop is a a framework desktop.
Guest 0
It's aware of all of those, so I can tell it to send an iMessage. Even though I'm talking to the Linux server in the cloud, it'll connect over open code to my, my my Mac Mac Studio to send the iMessage. Oh. So it's kinda like bring your host all the Node all devices you have to the open code server, and it kinda JS aware of all of them and all locations that exist. That's really cool because, like, everybody's talking about, like, oh, you have multiple panes talking to each other. You know? Oh, two Tmox panes talking to each other, or it opens the second Node, and, like, that's really cool. But, like, multiple
Wes Bos
machines being able to talk to each other is is really nifty.
Scott Tolinski
Yeah. I get that set up too as it's it's, like, endlessly productive
Guest 0
in terms of, like, management. Yeah. Yeah. It's funny because, like, we're we haven't even, like, lifted this into, like, a polished feature yet because the agent because all my devices are tail scale connected.
Guest 0
As long as the agent knows the names and the descriptions of each device, it'll just SSH into them and do stuff. So my my, remote server, when it wants to use the browser, it SSHs into my desktop and then uses that browser because that's where I'm logged in to everything.
Guest 0
So, like, it doesn't need anything special. Like, just the fact that they're connected is enough for it. Oh, cool.
Wes Bos
And Yeah. What's do you have any special, like, mobile setup? You know, we saw yesterday, Chris released the iOS app. Cloud's got their remote control thing.
Wes Bos
Any opinions on, like, what mobile should look like?
Guest 0
Yeah. We need to do a mobile app. It's been forever on our list, but we just it's just taking a while for us to get the core right Yeah. To support, like, this wide range of things that we wanna support.
Guest 0
Now that that's there, we'll probably kick off work on a mobile app. Like, we have, like, a shitty mobile web UI that's that works and Node use, but it's it's not it's not really great.
Guest 0
We just need to have clients everywhere, and they need to be good. Mhmm.
Scott Tolinski
Yeah. I actually think it's fine. I've been doing a lot of, like, what, Termeus on my phone to do open code or all that stuff.
Scott Tolinski
And I I find the open code GUI on mobile to be obviously the best experience so far. So it it just because that terminal stuff on iOS or mobile is just a giant pain in the ass right now still, or at least it is for me. So yeah. No. I appreciate you guys putting effort into having that GUI available at least for right now and Yeah. Whatever desktop app. I I yeah. Do you guys see the, official OpenClaw app? Woof.
Scott Tolinski
Oh, I didn't see it. No. I haven't seen it. It it just just came out, like, a day or two ago, and it is not good.
Scott Tolinski
Very yeah. It it just doesn't it's like somebody described it JS, I had to do a double take because I could not believe that this was the official one. It does feel very much just, like, tossed together. So, yeah, really appreciate the effort you guys put into your gooies in general.
Wes Bos
Can Can we talk about that? Because, like, I I wanna talk more about, like like, software engineering methodology, and, like, how you approach these things. Because, like, the the open co code, like, terminal app is significantly better than, like like, anything else that I've used. And, like, I switched over to the new Claude TUI, and, like, it doesn't even, like, scroll properly. And it's it's so frustrating. And, I I I would like to know, like, what what's your software engineering methodology that allows you to, like, have so much attention to the fit and finish of these products?
Guest 0
Yeah. I mean, I think we're all still trying to figure it out. Our team is struggling, I think, as much as everyone to Okay. To, like, balance all this stuff.
Guest 0
I think the first step is deciding that you care. I think that's the thing that I think it sounds obvious, but, like, there's so many rational reasons not to care. Right? Like, you'll see infinite arguments being, like, it doesn't matter that cloud code feels like this. It doesn't matter. They didn't try harder. They have, you know, billions of dollars in revenue.
Guest 0
So there's plenty of arguments out there to say that you don't actually have to care. You can kinda still be successful.
Guest 0
So one, like, do you actually care? I think our team does.
Guest 0
We have, like, other software we look at, and we're like, man, I wish we can make something as good as that. And that, like, that motivates us. So so that that's for one.
Guest 0
And then, again, I'm not saying you have to care. I'm just saying, I think for us, we still believe that it matters.
Guest 0
So that's where it starts. The second thing is, our token usage has, like, gotten pretty crazy now. Like, we use a lot of and this makes sense. Like, you know, we should, because we're experimenting with how all this stuff should be should be done.
Guest 0
This has changed recently. We've actually, as a team, in the past couple months, we're up five x in terms of monthly token usage. I'm not saying that to, like, be like, oh, we're so productive using so many tokens. I'm more talking about the, like, the models have achieved some kind of product market fit with our company to have that kind of growth in in a couple months. And and these Yarn with, like, some of the newer models that, are, like, you know, limited access still. But, the question is, like, what are we spending that on? Yeah.
Guest 0
What we're trying to do is really burn a lot of tokens on just kinda indulgently over designing everything. Like, if we're gonna do, like, even a simple API, an API to, like, read a file, what's, like, every possible way we can implement that API? Like, what are all the different prior art for anything other any other products that do this? What What are the different ways we can structure the responses? Right? So we'll spend a lot there, and this is something you can never do before. Like, maybe before you could think of, like, one or two ideas and you kinda go with your best one. Yeah. But we can be really indulgent there.
Guest 0
So that's, I think, very, very it's a good place to spend it because that actually results in better software, I think.
Guest 0
The second thing is Wes still think it's worth investing in primitives that maybe a coding agent can't just, naively one shot for you. Yeah. So a lot of reasons why our TUI is good is because we upfront invested in Open TUI, which is a, you know, TUI framework.
Guest 0
Now that's written in Zig. It takes a lot of meticulous effort by the people working on it to make sure it works across all different platforms. It's extremely performant.
Guest 0
They're constantly just, like, pushing that. And that is very coding agent assisted, but it's, like, very expert work. It's not the type of thing that the average person can do. It allows the average person like me to build something on top that JS, good, you know, and has a lot of capabilities.
Guest 0
But, yeah, I think for even with LLMs, like, you need Scott primitives to to build on top of. And, yeah, I think we think it's worth investing in that. We had the PR computer guys on, and they're working on
Wes Bos
on primitives as well. Right? Like, simple diffs, simple, sidebar trees, you know. And then, like, us idiots can take those beautifully designed primitives that smart people have implemented and then just slap them into our apps.
Guest 0
Yeah. There there's a million coding agent UIs now, and they all just use Peter. Yeah.
Wes Bos
Yeah. Including us. Five smart guys or not even five, like, two smart guys that built these things that the whole industry is built upon.
Scott Tolinski
Yeah. Yeah.
Scott Tolinski
Yeah.
Scott Tolinski
I'm wondering about something that we've been talking a lot on this show just to get your thoughts because you you seem to always have good takes on things. We've been talking a lot about, like, model routing Mhmm. And, like, viable techniques for routing to the correct model. Like, where do you think that's at? And do you think there is Yeah.
Guest 0
Things to evolve there? Yeah. I think this category is a little bit inflated because there's a whole set of middlemen that are desperately looking for something to do. Like, if you're not a model lab and you're like, I still wanna provide something compelling, and you're like, someone and and we're in that category. You know, we saw inference, and we're we're in the middle.
Guest 0
All you can really do is be like, well, what the model app can't do is use a different model because anthropic's never gonna serve you an OpenAI Node. And you're like, oh, we can do that.
Guest 0
Because it's something they can do, they kinda really talk about model routing. But sitting at that layer, I don't know how much there is you can do. Like, at best, when a request comes in, if the initial prompt is not hello, you know, sometimes I'm just like, hey, and then I send in then I send a bunch of stuff after. That initial prompt, like, maybe they can kinda figure out what what you should do.
Guest 0
But then after that, they can't really switch dynamically in the middle of a session because there's all kinds of, like, cost implications on that. Because if you switch models in the middle of the session, that new model JS like a full fresh cash bust, and it's like a very expensive switch. Yeah. So I think it's Yarn at that layer. The other side that we are very interested in, especially with some of the this, like, this, like, next generation of models, they are actually very good at this orchestrator pattern.
Guest 0
I think previous models I think people try this with previous models. I don't think they were good enough for the average person. But some people on our team, with some of these newer models, they've set it up where their primary session is this expensive model, but it's prompted to never actually do anything. It's prompted to only spawn sub agents for everything, and the sub agent is a cheaper model.
Guest 0
And this makes way more sense, I think, because the agent is kinda choosing, like, you can kinda granular use a cheaper model for, like, exploring, doing much of code changes, etcetera, but kinda still have the intelligence of the primary model. And I think in net, this ends up being cheaper, and it's also kinda like these newer models are very good at parallel work. So, like, you can, like, in a single session, be working on lots of different things with background sub agents that then kind of finish and and wake up the primary one. And it feels pretty good because you're just in a single session. Yeah. I think that's the Vercel time model routing that makes sense for me. You said a couple of minutes ago that, like, you're burning a lot of tokens and using some models that are not available yet. What are those? What do you what are you using? What do you got access to? It's very it's very unclear what I'm allowed to say and Wes I'm not. Like, I posted stuff that I thought was okay because I've seen other stuff posted. Probably people posted about it, and they told me it was not okay.
Guest 0
So just assume that we have access to, obviously, like, the you know, both of an anthropic have a pretty big, preview program to give to people for feedback. Mhmm.
Guest 0
So, you know, we we kinda see things a little bit early.
Guest 0
And the latest model from, I'm not gonna be specific, one of these labs is the one that caused us to go five x on on our usage.
Guest 0
So
Scott Tolinski
Really? So, Dex, who do we need to contact to get syntax, pre access to these?
Guest 0
I have no idea.
Wes Bos
He's not saying. But, like I don't believe. Let let me dig into that a second. One of these unreleased models has caused you to go five x on your token usage, and that is because
Guest 0
that's not because it's more token hungry. That's because it's changed the way that you're Our team isn't addicted to this thing. Yeah. And and, like, I wanna preface this by saying, like, you know, for people that aren't aware, we're a very conservative team. We've, like, for years, been pretty concerned about coding with AI. We're, like, very much not AI psychosis people.
Guest 0
We've been pretty measured about how we use it and what we how we communicate its capabilities.
Guest 0
But I will say, our team is is addicted to this this new generation of models. We actually just lost access because the preview period ended.
Guest 0
And for several days, all anyone would talk about is we're all, like, mourning the loss of this thing, and people are, like Yeah. What's the point of working anymore? Lots of AI generated funeral images, like, just, it's it's it's been tough.
Guest 0
Oh, man. So so I'm not even gonna say, like, it's it's because the new models are necessarily smarter. Like, it's not like they're like, oh, they're selling, like, so much smarter or they're, like, they kinda replace the whole human. It's not that. I think there's just there's just these, like, fine tuning things you can do around the usability of the model Wes they they can kind of I think they, like, kind of found a perfect zone where they, like, you can really trust them now. They kinda listen to what you're saying. They kinda pick up on stuff that you miss. Yeah. So it's not like it's, like, a human we're talking to all of a sudden. It's just a much better partner. And I think for our team Mhmm. You can see that in the numbers.
Wes Bos
What about models that are not from the big anthropic OpenAI, Google? You Node? Like, I know that you guys have, like, OpenCodeGo and and whatnot. Like, are those getting better? Like, where are we at with those?
Guest 0
Yeah. So now that we had lost access to this preview these preview models, we're now dropping back to, like, 5.5.
Guest 0
I think half our team dropped back to 5.5, but the other half are using GLM 5.2, which you've probably seen everyone talking about. Yeah. So I am using that.
Guest 0
Personally, I think it is very comparable to 5.5.
Guest 0
Again, after having used some of these newer ones, the the older ones all kinda feel roughly the same, so I'm kinda fine using anything. But they're definitely getting better. The fact that 5.2, to me, can replace 5.5 JS, sorry, GLM 5.2 can replace GPT 5.5.
Guest 0
Yeah. I mean, they're they're they're getting better, and they're just gonna keep getting better. And I I personally find that gap is closing more and more. I think the frontier models will always have some advantage, just because they it's one of those things where whoever starts first kind of always stays ahead, because there's, like, some compounding factors there. Mhmm.
Guest 0
But, man, like, Wes see a ton of usage on Go.
Guest 0
People use it entirely for all their work.
Guest 0
Really? Yeah. Like, it's again, I think we're maybe in a bubble that has really high the salaries are high.
Guest 0
The value of your currency is really high. We have access to kinda spending lots of money on these frontier models. But for most of the world, that's not the case, even within The US.
Guest 0
Like, when we launched Go, which is our cheap plan for meant for, like, open source models, we were thinking, oh, it's like an international plan for people across the world. But The US JS I think it's still our number one subscriber. Like, most of our subscribers in The US not most, but it's, like, in the ranking, it's it's number one.
Guest 0
So yeah. Like, the world of developers and people that wanna code is very, very large, and not all of them are even the $200 a month plan is is out of reach for a lot of
Wes Bos
them. Yeah. Yeah. That, like, pricing, I'm curious, like, what you see the future of that being, the people that are just running this stuff all day long. Are we gonna are we gonna get to a spot where companies are spending
Guest 0
thousand bucks a month, $2,000 a month per employee? Or do you think that this the prices of the stuff is gonna stabilize as new chips or whatever come out? So we we've been doing some math again because of our cup. So we're looking at last month of data as our company has kinda exploded in usage, and we calculated how much that cost.
Guest 0
And we look at it relative to payroll. And for us, like, you know, this is, like, a lot of usage for us. Like I said, it's five times more. Yeah. It's roughly, like, 15% of our payroll. So if you take about the amount we're paying our team, it's, like, another 15% tax on top of that to give them access to to these models.
Guest 0
That's, like, not that bad. You Node, again, for a company like ours in tech, you know, like, we typically make a lot of revenue per per employee. Like, 50% is, like, very negligible in the grand scheme of things. Not the case for every industry.
Guest 0
But also these prices are gonna come down can come down a lot. Like, if you are price sensitive, these open source models are just way cheaper.
Guest 0
I think this is confusing to people because there's a lot of headlines about open AI and anthropic just lose money and, like, they're never gonna be successful.
Guest 0
But margins on inference are, like, insane right now, especially because OpenAI and Anthropic are increasing prices.
Guest 0
I'm estimating they probably make 90% margin on their sales, which means yeah. Which means breakeven could be 10 x cheaper.
Wes Bos
Man.
Wes Bos
Somebody told me that. They said there's a 70% markup on inference. Scott like training or anything like that, but that's good to hear. So 90%.
Wes Bos
That obviously isn't the cost of of training these models. Right? Mhmm.
Guest 0
Yeah. Of course. So there there's, like, r and d, but, you know, like, as a business, you separate those things out separately because you can shut off r and d and still make still make money.
Wes Bos
Yeah. And what about the crazy people that think you can run it locally, including us? What's your take on the people that think they're gonna have a machine running in their in their
Guest 0
Yarn? I'm very careful about talking about this because this community gets very angry.
Guest 0
So I'll again, I'll preface this by saying, there's a lot of good reasons that people wanna run stuff locally. If you just do not want stuff outside of your house, makes total sense. You can kinda go down this path.
Guest 0
But if you think about the underlying and if if it's a but if you focus on cost, it's not really like, the local model isn't gonna help you on cost. Because any mechanism that makes it cheaper to host locally makes it, like, 10 x cheaper to host in the cloud. Like, if a model gets more efficient or more capable at a smaller size, that's just gonna be cheaper per token, in in the cloud. So I think local model is more of a privacy thing, less so a a cost thing.
Guest 0
And just to kinda give you guys a fewer numbers. So, like, again, because we are an inference provider, we see some of these details.
Guest 0
We still use middlemen.
Guest 0
Despite using middlemen's for hosting our GPU's, there are some models that we are able to host at a 70% discount to us.
Guest 0
That is very, very cheap, which means we can make 70% margin by selling it at sticker price.
Guest 0
And that's with a middle man involved. So if you directly spend the capital to acquire the GPUs, you can probably hit those, like, 90% margins that I'm estimating for Anthropic.
Guest 0
Which means, like, Scott inference is is very, very cheap. Yeah. Again, these are for open source models. Yeah. So we we still have to rely on those getting better.
Guest 0
But, you know, it's going in that direction so far. Yeah. That's encouraging.
Scott Tolinski
Let's talk about cloud code.
Scott Tolinski
So Cloud Code, it seems like it's been very unclear of what their stance is, whether or not even providers like OpenCode can use things like the Cloud Code Max plan.
Scott Tolinski
Or it it it just seems like they're probably a difficult partner to work with. But what is the current status of Cloud Code, the Cloud Code Max plan, OpenCode, and third party harnesses?
Guest 0
Yeah. So, the integration that the plug in that we had in OpenCode that let you use your max plan, that's definitely not allowed.
Guest 0
They we fought with them on that for a lot, and we did not win. So, that that that's definitely not allowed. Of course, people still find ways to hack it in. We just can't officially support it.
Guest 0
The, the SDK, which is, like, spawning Claude or, like, using Claude headlessly, that is now in a gray area. They're saying, for now, it's allowed. So a product like Conductor can wrap it. A product like t three code can wrap it. We're never gonna wrap that. Like, that just kind of defeats the point about code in a lot of ways.
Guest 0
So the the orchestrator thing JS, like, the the cusp or, like, the alternative UIs for these things. I think those products work for now. But, again, still unclear.
Guest 0
It this is Node of those things where this comes down to your company's culture. Like, fundamentally, whether you're a very consumer oriented business or you're a, enterprise oriented business.
Guest 0
OpenAI is very much a consumer oriented business, which means they will burn any amount of money, raise any amount of money to bring the experience to more people, which is why the OpenAI subscription is supported in Node officially.
Guest 0
I think Anthropic doesn't really have that exact same culture. I'm not saying one or for the right over the other. It's just, like, different ways that a company can operate.
Guest 0
If your company JS not oriented like that, it's hard because any inference you're allocating for this, like, more consumer oriented thing, there's a sales guy being, like, I've got an enterprise customer willing to pay, like, actual prices for this thing.
Guest 0
If your compute is all at all limited, it becomes very hard within your organization to to justify allowing open code users to use this thing.
Wes Bos
And I I think they have more compute Node, but yeah. Is that why they're they don't want you using it? Because, like, everyone's like, what does it matter? I'm paying you for my subscription. Why does it matter where I'm using it? And I hear people saying like, oh, they want the data for training.
Wes Bos
They want control over it. But it simply is just like there's they're compute bound.
Guest 0
Well, the the the biz every business is is a form of a funnel. Right? You have some of the top of your funnel that draws people in, and you ideally get them all the way down.
Guest 0
They're oriented around cloud code being top of funnel. It's a very consumer oriented product.
Guest 0
It gets in, you kinda get your company to start using it, and then your company starts paying per token prices.
Guest 0
That might not necessarily happen if they're using something like Node instead.
Guest 0
They're kind of, like, letting us to siphon get into their funnel.
Guest 0
And people might not end up converting all the way down in that way because you can switch other models in open code. Yeah. So if you're not liking Claude, like, you can just switch to whatever the newest, hottest thing is.
Guest 0
So I think that that's one thing. And the second thing again, it's, like, there's competing yeah. There's competing demand for the for the compute. Like, anything you invest on the top of the funnel, you'd better be able to make a case that's gonna come back down to the bottom. Mhmm.
Guest 0
And, again, if you're consumer oriented, you're a little bit more flexible on this.
Wes Bos
Do you ever see a world where there's a new model released and you simply are not there's no API for it? You can only use it through their app? Like, even just, like, I look at, like, Eleven Labs. They have a sick app, but you can't use it unless they they do have an API, but you can't use it unless you, like, subscribe to their, like, monthly thing. You Sanity pay, like, 20¢ for just one thing that you need. You think that's coming?
Guest 0
Yeah. And I think this JS, again, like, it's this is a reflection of the company's internal structure. So if you look at product teams, product teams are gonna be very pro this because they're like, we can create a very specialized model. We can create a very specialized product around it. We can trap it together so that if you wanna use a model, you have to use a product. That's a very good setup. You know, it's a go it's a very good lock in setup for a product oriented team. But if your sales organization has revenue goals, they're like, okay, our revenue goals are 100,000,000,000.
Guest 0
The best case scenario of your API or your product only model is 50,000,000,000, and we have to fill that 50,000,000,000 gap. The sales team is gonna be like, Node, we gotta put this in API, because we can we can hit our goals better that way. So as long as that dynamic is there, it's gonna be very hard for our organization to justify, like they're basically giving up some revenue to try to lock in more market share. Mhmm. I'm not saying they can't ever make that argument. They might be able to. I am very concerned about that. I think as these labs move more into a product layer, they kinda have this unfair button they can press, and I wouldn't be surprised if they press it at some point.
Guest 0
And and and they'll justify it in weird ways. They'll be like, this model is really dangerous, so it's only safe to use it within our harness. We can't let people use it in other harnesses, which isn't the actual reason, but that's likely how they'll frame it. Okay.
Wes Bos
Yeah.
Wes Bos
One one more question about that safety stuff is, like, Fable, unsafe. All of these new models, unsafe government.
Guest 0
Is that true, or is that just are they just hyping it up? I think there's, like, a lot of things that are true, and they're maybe somewhat in conflict with each other. These models do have potential for great harm. I think it's okay for the government to say Wes need to have some kind of review process before we put this out there, so we can kinda adequately, you know, regulate this. Yeah.
Guest 0
This isn't that weird. If you're at a big company like Meta, when they put out a product and it has, like, an upload image feature for, like, their profile picture, they have to show the government that they're doing, like, child pornography filtering on that. And it's, like, crazy amount of regulation at that scale. So, like, over the most trivial features in the app.
Guest 0
So I get that there Node to be something.
Guest 0
It would be really bad if this process is very ignorant or corrupt, where the end result isn't blanket approval of the model for broad access for everyone. Mhmm. If the end result is unequal access tied with the government or a government process, that's just gonna a really bad situation. And I hope that that's not the case. I hope it's a more boring outcome Wes they they just need to do some kind of process that takes a month every time they release a new model.
Guest 0
That's what it's looking like so far, but you never know. The flip side is there this isn't, like, a hyper rational thing.
Guest 0
I think it's not a good idea for these labs to start making crazy claims that they have a nuclear weapon. Because it's it's gonna attract political interests. I mean, you attract political interests, that's not always a very rational process.
Guest 0
It's like you're playing with a bomb there. It could end up in a boring situation, or it can be something really bad, like, the the regulation is is incorrect and overly aggressive and bad for everyone in the economy. So I I wish these labs would be a lot would be more careful in terms of public perception, because you can't just go around saying you have a bomb and expect nothing to happen. You know? Yeah.
Scott Tolinski
Yeah. Let's talk, MCP skills, just general tools Mhmm. When coding with AI. What's worth actually paying attention to and using today? What are you all,
Guest 0
maxing on? Yeah. So I think most of us use pretty vanilla setups.
Guest 0
We'll have what's interesting is our Discord bot and this is this is, like, very similar to the Slack tag product they had.
Guest 0
So WebNCODE runs as a Discord bot. This is only one we'll make public at some point.
Guest 0
This thing has a lot more MCPs and skills than our than our personal setup.
Guest 0
We have this fun one that, Kit Langdon came up with called Gang Growth.
Guest 0
Basically, whenever we're stuck in any kind of design question for well, it could be literally about the business. It could be about API design. It could be about, like, implementation.
Guest 0
We'll kinda describe the product. Again, we use voice prompting for everything. But in the Discord, we'll tag open open code. So, like, it's like a nice skill to, like, get our team to collaborate together on stuff, which which is pretty fun.
Guest 0
The Discord bot also has, like, a ton of MCB servers and tools that, just connect to all company systems. So, like, it has access to our full data lake, which means I can ask you fairly complex Wes, like, in the past week, which go of all the Go subscribers that opted into extra billing, how much do they spend in total? And it'll go, like, kinda kinda figure that out. There's basically you know, we're talking about this yesterday. Like, there's no reason to tag another person.
Guest 0
You basically always if you have a question or you need something done, you tag OpenCode first. And if the other person sees it, they'll jump in, and they'll kinda you guys will all work together on it. But, Open Go often often can solve it. So that's where you kinda overload a lot of tools and a lot of, stuff. We have, like, code mode implemented there, to make that all efficient. But, yeah, on our personal setups, I don't think any of us do do too much.
Wes Bos
And that Bos, So if that bot needs to then bring you back information or whatever. Right? There's there's this new talk of, like not necessarily new, but, like, the MCP UI or MCP apps or or WebMCP or, simply just spinning up an HTML file and linking you to it.
Wes Bos
Mhmm. What's what do you see as the future of, like, your coding agent needs to show you something that can't simply just be shown in a terminal? Yeah. So we're definitely gonna add some kind of artifacts feature, into open code so that it can
Guest 0
produce a document and send it to you. Because it's it's very cool when it, like, writes up, and it uses HTML to, like, visualize something for you with SVG or something. I think that that's pretty awesome.
Guest 0
And that doesn't require anything special. That's just, like, a way of using the the the agent. Mhmm. I haven't looked too hard at MCP UI. I think it is something we'll support in our desktop app, especially when we start to focus on non technical people.
Guest 0
Because I think their the questions they ask and the things they need to do probably benefit from some dynamic UI or something something a bit richer.
Wes Bos
I haven't looked too deeply into these specs to know how how they work or anything. But The dynamic part of the spec is not there yet, so it's probably not worth worth doing just yet. Because right now, it's just like, here's something, and then, like, update the number of limes in my shopping cart, and then it gives you an entire new it's slow. It's it's it's I'm I'm bullish on it, but it's it's still not there quite yet.
Guest 0
The other thing is I don't know how much this part has taken over, but another thing our team's addicted to now is voice prompting.
Guest 0
Yeah. To the point where we don't even use even when we're just messaging each other on Discord, we're all just using voice prompting because we just hate typing now.
Guest 0
And when you can just talk at it, a lot of UI, especially interactive UI, JS, like, I'd rather just describe roughly what I need to do. If I had to type it all out, that would suck for sure. But given I can do voice, and voice is so fast now, and it runs locally, it's it's just so good.
Guest 0
I'm very happy just voicing everything.
Scott Tolinski
What app do you use for voice? Vercel tip, I got, a foot pedal because I've been doing voice prompting. Nice. And I gotta say, push I have one that's, like, push to enter and the other one that is, like, push to to trigger, dictation.
Scott Tolinski
And then this other one changes my tab in which Asian tab I'm going on. But I'm just sitting back, you had, just chatting and and prompting of this thing. Yeah. It's kinda sick. Yeah. Whenever I bring this up, I get a lot of people being, like, a little skeptical about it. And I totally get it because
Guest 0
I wasn't doing it until I saw, like, someone doing it. Like, like, Kit was the one that brought to our team.
Guest 0
We saw him doing it, and, like, the moment you see someone doing it, it, like, unlock something in your head. Because, like, I think Wes you're not doing it, you imagine it to be this kinda weird awkward thing. But it's the most natural thing Vercel you can ramble. You can be confusing. You can, like, mess up words. It can, like, kinda mess up the dictation. It all doesn't matter, because it's what the l m is good at. Mhmm. Understanding what you actually meant.
Guest 0
Yeah. And you're are you using hex from Kit? Obviously, that's the app you use. Well, so I I use Linux. So I I use Hex on my Mac, but, on my main machine, I just use Handy.
Guest 0
Yeah. Yeah. I think there could again, I think I think Hex is great on Mac. On Linux, it feels, like, a little bit rough.
Wes Bos
But, you know, it's the models are very good. This kinda what matters. You just have it hooked up to a a key, you know, like a keyboard shortcut, or how do you trigger it? I mean, now that I saw the foot pedals, that's a great idea. Man. Maybe I should I should do that too. I have it I have it hooked up to the little button on my mouse Wes I double tap it. That's clever. And then somebody has sent me a DM yesterday. They made a ring that you tap Yes. Which is kinda cool. They're gonna send me one. I'll try it.
Guest 0
Yeah. The, the pebble guy is doing a ring with a button on it. And that's for, like, a different use case, but Mhmm. I liked it because it's, like, a subtle gesture. Like, it's basically, like, you know, it's kind of always there.
Guest 0
I I need to try it out to see if if it actually works. But Yeah. Yeah. Right now, I just my fingers are still on my keyboard most of the time anyways, so I just have a a key binding there. Yeah. Sweet.
Wes Bos
Anything we haven't touched on that you wanna touch on? Any opinions or whatever that you have?
Guest 0
Not really. I'm just excited for this next generation of models. I think, mostly Wes a new model comes out, I it's, like, basically feels the same. You know? And I usually post it may make a post making fun of that fact.
Guest 0
But this is the first time where I feel like I think these might click a lot more for for a lot of people. So I'm pretty excited for again, they're technically released. They're just Scott, you know, they're not the government's not allowing people to use them. Yeah. Us, plebs.
Wes Bos
That's a question I have JS testing these things. You Node, there's all these, like, benchmarks and whatever, and then you have all these, like, Scott.
Wes Bos
And then you also have just, like, people's, like, vibe of, like, it feels a lot better.
Wes Bos
Do you think we'll ever get to, like, a a a benchmark that makes a lot of sense?
Guest 0
Yeah. At this point, I don't think I look at benchmarks at all. I don't know if I ever really did. I don't think anyone ever really did. I think they got they just kind of became white noise at some point. Yeah. We all know that the numbers go up. Like, congratulations number went up. You know? Yeah.
Guest 0
And it's like, it it Wes up more than the other guy, but then when the other guy releases, their number goes up more. So it it's it's I don't know if that means anything to us.
Guest 0
So I just look at I mean, I love qualitative feedback. I love people being like, sharing what they've been able to do or what they've built. We have you know, you obviously don't get that at, like, a million data ESLint Scott, but, these are, like, products at the end of the day, and they're kind of fuzzy.
Guest 0
And it comes down to are people happy, are people frustrated? That's why I like looking at token usage on our team.
Guest 0
Just this because it because to me, that's, like, our if we if I see it go up, that means something is working. Like, they're they're liking something.
Guest 0
Yeah. So yeah. I think I I still rely on a lot of qualitative feedback, and people kinda demonstrating how the how they're using it.
Guest 0
Given our team is you know, within our team, we have cloud fans. We have GPT fans. We've got some open source model fans. So we have okay coverage on visibility and until that.
Scott Tolinski
Yeah. So, Dax, this is your second time on the show, so you know the drill. We have sick picks and shameless plugs on this show. Sick pick being anything that you're just enjoying in life right now. Do you have something that you would like to share as a sick pick? Yeah. I mean, it's gotta be that thing I mentioned earlier,
Guest 0
e x e.dev.
Guest 0
I think, like, yeah, if you wanna experiment with this machine in the cloud thing, it's a really clever product. It's put together really well. Like, if you use Tailscale, you you know how good Tailscale is. Like, Tailscale just, like, freaking works. Mhmm.
Guest 0
And this has, like, that same vibe to it. And and and and I really love, like, businesses that Yarn, like, I think a lot about positioning.
Guest 0
And this product is it's, like, positioned in, like, a weird vacuum that existed. Like, you can obviously rent servers, whether it's from AWS or for whatever, But you couldn't really easily rent a server with a fast disk that is persistent, that is at a good price. So it, like, feels like there's really deep gap that was only filled before by, like, shady VPS providers that you kind of find and disappear.
Guest 0
No joke. This is like a serious thing. I had posted about this a few years back.
Guest 0
Back when my first iteration of my dev box, I was just trying to find the cheapest possible thing, and I found this VPS provider in Miami.
Guest 0
He literally faked his own death, and the machine went offline.
Guest 0
He post he sent an email to everyone being, like, I'm going in for a medical procedure, so I'll be unavailable for three days. And then my server goes down three days later. So I'm, like, oh, Node. Did something happen? And it's been a month, and no no one's heard from him. Eventually, I find a forum post. People kinda track this guy down, and he had a previous VPS service also that also disappeared under similar circumstances.
Guest 0
So, yeah, that whole market for, like, cheap performant things, it was just, like, not very reliable. I don't even get what the scam was. Like, I was paying him money for a service. Like, why did he need to, like, disappear? Like, I I don't fully get it still to this day, but
Wes Bos
yeah, I don't know. Wow. That's that's incredible. Cool. And the last thing, shameless plugs, would you like to plug to the audience?
Guest 0
Oh, my own stuff.
Guest 0
Man, I don't know. I I guess open TUI. You Node, if you're, building a TUI, it's a great way to do it. Can you build a nice performance two e with React or solid JS? Or I think we have maybe view bindings. I don't remember.
Guest 0
It's what OpenCo is built with.
Guest 0
We're working towards a one point o on that soon.
Guest 0
Yeah. It's been a nice renaissance of terminal products, apps.
Guest 0
And it's a great way to JS built on this? Is the the new grok build or the x AI stuff built on this? The grok build CLI is very well done. Like, it is a very well executed for performing CLI, but that that's that one's straight in Rust. I think they might be using, this library called Ratatouille.
Guest 0
Okay.
Guest 0
So that one is not.
Guest 0
I think the new Hermes agent TUI is built on, on Open TUI.
Guest 0
There's been a few others. I I can't remember off top of my head, but, it's growing quite a bit. And, like, pretty much whenever I see a TUI on my timeline nowadays, it's usually with Open TUI. Sec. Especially because you can you can vibe code it pretty well because it's just React.
Wes Bos
Yeah. Beautiful.
Wes Bos
Alright. Wes, thank you so much for coming on. Appreciate all your time and all your insights. This is gonna be a good one.