The lower-post-volume people behind the software in Debian. (List of feeds.)
Hvorfor it-afdelingens budget ofte rammes af valutakursudsving
Globale it-investeringer og driftsomkostninger er i vid udstrækning underlagt internationale markedskræfter, hvilket gør it-budgetter særligt sårbare over for fluktuationer på valutamarkedet. Ifølge en analyse fra analysehuset Gartner i 2026 udgør software- og cloud-tjenester nu over 35 % af de samlede globale it-omkostninger, og størstedelen af disse afregnes i amerikanske dollars (USD). Når den danske krone (DKK) svækkes over for dollaren, stiger de reelle driftsomkostninger for danske virksomheder øjeblikkeligt, selvom det faktiske forbrug af ressourcer forbliver uændret.
Mange virksomheder opererer med faste årlige budgetter, som lægges i fjerde kvartal af det foregående regnskabsår. Hvis en virksomhed har budgetteret med en USD/DKK-kurs på 6,80, og kursen i løbet af året stiger til 7,20, repræsenterer dette en stigning i omkostningerne på knap 6 %. For en mellemstor it-afdeling med et årligt cloud-budget på 500.000 USD svarer dette til en uforudset ekstraudgift på 200.000 DKK, hvilket kan tvinge ledelsen til at udskyde andre kritiske udviklingsprojekter.
Særligt hardwareindkøb påvirkes også af disse mekanismer, da mikrochips, servere og netværksudstyr produceres globalt med afregning i USD. Selvom udstyret købes gennem en dansk distributør, er priserne ofte indeksreguleret i forhold til dollarkursen på leveringstidspunktet. Dette skaber en uforudsigelighed i forsyningskæden, som kræver konstant overvågning og finansiel risikostyring.
“Valutarisiko i it-afdelingen er ikke længere et perifert problem for finansafdelingen. Det er en direkte operationel udfordring, som kræver præcise værktøjer og proaktiv risikostyring for at undgå budgetskred.”
– Finansiel analytiker i it-sektoren, 2026
SaaS-licenser og cloud-hosting: De skjulte valutaomkostninger i USD
Software as a Service (SaaS) og Infrastructure as a Service (IaaS) har ændret it-afdelingens omkostningsstruktur fra kapitalomkostninger (CAPEX) til operationelle omkostninger (OPEX). Store udbydere som Amazon Web Services (AWS), Microsoft Azure og Google Cloud Platform afregner som standard deres ydelser i USD eller EUR. For europæiske virksomheder betyder det, at den månedlige hostingregning ændrer sig i takt med de globale valutakurser.
Mange af disse cloud-baserede services kører i øvrigt på open-source infrastruktur, hvilket du kan læse mere om i artiklen Hvad er Linux? – En komplet guide til det åbne styresystem. Selvom selve styresystemet ofte er licensfrit, betales der stadig store summer for den underliggende hardwarekapacitet og supportaftaler, som næsten altid afregnes i udenlandsk valuta.
For at styre disse fluktuerende omkostninger anvender mange finansansvarlige en digital valutaberegner, som henter realtidsdata direkte fra de globale finansmarkeder. Dette gør det muligt at estimere de præcise månedlige udgifter i danske kroner og foretage de nødvendige budgetjusteringer i tide. For præcise, bank-verificerede kurser kan man også anvende en ekstern Valutaberegner – omregn opdaterede valutakurser til at verificere de daglige transaktioner og sikre mod ubehagelige overraskelser på kreditkortudtogene.
Udover selve hostingafgifterne er mange enterprise-softwarelicenser (såsom Salesforce, Adobe Creative Cloud og Jira) bundet til dollarpriser. Selv når faktureringen sker via en europæisk enhed i euro, er prisen ofte fastsat ud fra en omregningskurs fra USD, hvilket betyder, at valutasvingninger indirekte overføres til den europæiske køber.
Styr på økonomien ved outsourcing af it-udvikling til udlandet
Outsourcing af softwareudvikling til lande i Østeuropa eller Asien er en udbredt strategi for at reducere lønomkostninger og imødekomme manglen på kvalificeret arbejdskraft i Danmark. Men når man samarbejder med eksterne udviklingshuse i eksempelvis Polen, Ukraine eller Indien, introduceres der komplekse valutarisici. Kontrakter indgås ofte i lokal valuta (såsom polske zloty (PLN) eller indiske rupees (INR)) eller i USD for at skabe en fælles standard.
Hvis den lokale valuta i modtagerlandet styrkes over for den danske krone, stiger timeprisen for udviklerne tilsvarende. En stigning på 10 % i den polske zloty kan hurtigt eliminere den økonomiske fordel, der oprindeligt var ved at outsource opgaven frem for at løse den internt eller med lokale konsulenter. Det er derfor afgørende at indbygge valutaklausuler i samarbejdsaftalerne eller foretage finansiel kurssikring (hedging).
Når der afregnes med udenlandske leverandører, skal moms- og skatteforhold indberettes korrekt til de danske myndigheder, hvor man med fordel kan konsultere SKATs officielle Valutaomregner – info.skat.dk. Korrekt omregning på faktureringstidspunktet er afgørende for at undgå efterreguleringer og sikre, at bogføringen overholder gældende dansk lovgivning.
Brug en præcis valutaberegner til at sikre dine it-budgetter
For at imødegå de risici, som valutafluktuationer medfører, bør it-afdelingen og økonomifunktionen samarbejde tæt om at overvåge markedet. En manuel omregning baseret på tilfældige søgninger på internettet er sjældent tilstrækkelig, da kurserne ændrer sig sekund for sekund. Anvendelsen af professionelle værktøjer sikrer, at man altid arbejder med de mest aktuelle og præcise tal.
Der findes forskellige finansielle strategier til at minimere valutarisikoen, og valget afhænger af virksomhedens risikoprofil og budgetstørrelse. Nedenstående tabel illustrerer de mest almindelige metoder til håndtering af valutarisiko i forbindelse med it-indkøb:
| Strategi | Beskrivelse | Fordele | Ulemper |
|---|---|---|---|
| Spot-kontrakter | Køb af valuta til den aktuelle dagspris. | Enkel administration, ingen langsigtede forpligtelser. | Ingen beskyttelse mod fremtidige kursstigninger. |
| Terminskontrakter (Forwards) | Låsning af en valutakurs til levering på en bestemt fremtidig dato. | 100 % budgetgaranti og fuld økonomisk forudsigelighed. | Man går glip af potentielle gevinster, hvis kursen falder. |
| Valutaoptioner | Retten, men ikke pligten, til at købe valuta til en fastsat kurs. | Maksimal fleksibilitet og beskyttelse mod tab. | Kræver betaling af en præmie (gebyr) up-front. |
Ved systematisk at anvende en præcis beregner kan it-chefen hurtigt simulere forskellige scenarier. Hvis dollarkursen eksempelvis stiger med 5 %, 10 % eller 15 %, kan man øjeblikkeligt se konsekvenserne for det samlede årsbudget og træffe beslutning om, hvorvidt der skal foretages kurssikring via virksomhedens bankforbindelse.
Sådan integrerer du valutadata direkte i dine egne it-systemer
For større virksomheder er manuel indtastning af valutakurser i ERP-systemer (Enterprise Resource Planning) både tidskrævende og forbundet med risiko for menneskelige fejl. Den mest effektive løsning er at automatisere processen ved at integrere realtids-valutadata direkte i virksomhedens økonomisystemer via et API (Application Programming Interface).
Moderne ERP-systemer som Microsoft Dynamics 365, SAP eller NetSuite tilbyder standardmoduler til automatisk valutaopdatering. Ved at forbinde systemet til en pålidelig datakilde via en REST-API kan systemet hente de officielle lukkekurser hver nat. Dette sikrer, at alle indgående fakturaer i fremmed valuta automatisk bogføres med den korrekte omregningskurs på transaktionsdatoen.
En typisk integration fungerer ved, at et script (eksempelvis skrevet i Python eller afviklet som en serverless funktion i cloud’en) kalder API’et, modtager data i JSON-format og opdaterer databasen. Dette reducerer tidsforbruget i finansafdelingen markant og minimerer risikoen for differencer i årsregnskabet, som skyldes forældede eller fejlagtige kursberegninger.
Ofte stillede spørgsmål
Hvorfor svinger it-budgetter på grund af valutakurser?
Mange it-ydelser, herunder cloud-hosting (AWS, Azure) og softwarelicenser (SaaS), afregnes i amerikanske dollars (USD). Når den danske krone svækkes over for dollaren, stiger de reelle udgifter i danske kroner, selvom det faktiske forbrug af it-ressourcer forbliver uændret.
Hvordan påvirker USD-kursen prisen på cloud-hosting?
Cloud-leverandører prissætter som regel deres globale infrastruktur i USD. Hvis dollarkursen stiger med f.eks. 8 % i løbet af et regnskabsår, vil den månedlige regning for cloud-hosting stige med nøjagtig samme procentsats for danske virksomheder, medmindre der er indgået en fast prisaftale.
Hvad er forskellen på en standard valutaberegner og en bank-kurs?
En standard valutaberegner viser typisk midterkursen (spotkursen) på det globale interbank-marked uden tillæg. Når man veksler penge eller betaler fakturaer gennem en erhvervsbank, pålægger banken normalt et vekselgebyr eller et kurstillæg, hvilket gør den reelle afregningskurs en smule dyrere.
Kan man automatisere valutaomregning i sit ERP-system?
Ja, de fleste moderne ERP- og økonomisystemer understøtter automatisk integration via API’er. Ved at opsætte en daglig synkronisering kan systemet automatisk hente og opdatere valutakurserne, hvilket eliminerer manuelt arbejde og sikrer præcis bogføring af udenlandske fakturaer.
While you (yes, you! no, not you, the one behind you) have been sweltering in the heatwaves of the northern hemispheres (Assisted-by: AI), I've been busy adding graphics tablet support to libei. This is scheduled for the soon to be released libei 1.7.0.
The initial work was done by Jason Gerecke and Josh Dickens from Wacom, I've been extending, polishing and testing it for the last few weeks.
Also, upfront: this only covers the stylus part of a tablet, we do not yet have an implementation for the "pad" part (the buttons, dials, rings, strips).
libei is, of course, the library for Emulated Input, a good-enough transport layer for sending logical input events between processes. We're already using libei as part of the XDG Portal Remote Desktop and Input Capture portals where we've been busy hurtling key and pointer events between the participating parties (and soon gesture events and text).
In the next release of libei, we will now also have "ei stylus" capabilities, i.e. the ability to send tablet stylus events. Getting pointer, keyboard and touch events supported was a long undertaking, everything was new and shiny and needed to be added everywhere in the stack. Now that all this is in place, scuffed and scratched, adding tablet events will be quite simple.
The ei stylus interface
Here's a short outline of how libei handles tablet events because it is, of course, different to how libinput handles them. Logical events are much nicer after all than physical hardware events.
First: we have a new interface: "ei_stylus". An EIS implementation (e.g. your compositor) may provide you, the libei client, with a device that supports this interface and one or more associated regions (typically representing the available screen areas). Typically this will be a separate device to the pointer devices or the keyboard devices but it's not a requirement. The ei_stylus interface comes with a bunch of capabilities you'd expect from a stylus (tilt, pressure, distance, ...) that you can selectively enable to emulate the stylus you want to. So basically, EIS will say "here's a stylus device, I support pressure, tilt, rotation, ..." and then the libei client says "This stylus should have pressure and tilt but nothing else". And then you do the normal thing: send proximity events, send tip down/up events, send data for the various capabilities you've enabled.
Happily for the EIS implementation, libei forces the client to take the guesswork out of everything: if you select the pressure capability, you must send a pressure value when coming into proximity. Where libei is used to forward data from a physical stylus (e.g. via some remoting protocol) it is up to the client to deal with firmware bugs that e.g. won't send data until a few frames in.
Note that there is no "tablet" anywhere. The tablet is represented by the region that the device may interact with. So in some ways every tablet is an on-screen tablet (which makes sense since we have logical events).
Multiple styli
The only quirky thing is how to request multiple styli[1]: libei 1.5.0 has added a "request device" request that allows a client to say "hey, EIS, I want a new device with capabilities pointer, keyboard, ...". And, if you've been a nice client, minding your own business, the EIS implementation may just create such a device for you.
So for the case of multiple styli: if the default stylus (if any) isn't good enough, you can now tell EIS that you want a(nother) device with stylus capability, configure the stylus capabilities once the device shows up and voila, you now have a normal pen, an art pen and maybe even an airbrush represented as logical device in libei. And since they're all separate devices in the protocol, they can be individually tracked and used, much like libinput tracks individual styli.
[1] For the "lots" of users that actually use multiple styli...
/me gestures vaguely at everything
Oh, hey, this works now? Great!
libei 1.7.0 (to be released soon) comes with a new interface: "ei_gestures" which, creatively, will allow for gestures to be sent between a libei client and an EIS implementation (typically: a Wayland compositor).I'm not going to go too deeply into how pinch, swipe and hold gestures work, suffice to say we've had those in libinput (for touchpads) for years now so compositors and toolkits should already support those. And since libei and libinput have vaguely equivalent API layers integrating gestures for libei devices in compositors should be fairly straightforward.
The plumbing layers in the portals exist already too, so adding gestures to libei means that - once the compositors support it - we can have gestures support in remote desktop and input capture implementations without needing to update anything else. Hooray! Join in with me. Hooray! Louder! HOORAY!
For testing I had a (vibe-coded and thus immediately abandoned once testing was complete) gesturemouse utility which translates input events from a mouse into gesture events (depending which button is down). But don't let my lack of be a limit to your imagination, I'm sure you can come up with good use-cases for this.
If you've been paying attention (and I know you have, because it'd be embarrassing for you if you didn't) you'd have noticed that libei 1.6 (May 2026) added support for keysym and text events.
libei sends logical events between a libei client and an EIS implementation (typically: a Wayland compositor) but the keyboard interface it had was designed like real keyboards: key codes together with an (XKB) key map. You press one key, the keymap decides what that key means on the compositor side and off we go. This is easy but not always useful.
As of 1.6.0 libei now also supports an "ei_text" interface. A compositor may choose to provide you[0] with a device that supports this interface and that gives you two really nice opportunities.
First, you can now send a key sym. Instead of sending the KEY_Q
key code and hoping it actually translates to 'q' (and if there's e.g. a
frenchman^Wfrenchperson lurking behind the keyboard it may mean 'a'), you can
now send 'q' as actual keysym. Or 'Q' instead of sending shift+q and
hoping for no french influence in the process. It becomes the EIS
implementation's job to handle that keysym - if it's a shortcut it may handle
it directly, otherwise it may pass it on via Wayland to an application[1]. This
centralises the keysym to keycode handling in the EIS implementation which is a
pain for compositor authors (though they likely have that code already for e.g.
RDP support) but reduces the variety of differently-wrong implementations in
clients and of course makes it so much simpler to write clients.
Second, a client can send UTF-8 text to the compositor. So instead of emulating shift, keycodes, etc. you can literally send "Hello World" and expect the EIS implementation to pass that one. Again, makes a bunch of utilities a lot simpler to write and I mostly leave it up to your imagination to figure out what to do with that.
Notably for both cases: libei is about logical events that have a specific meaning that do not need further interpretation. If a client sends 'Q' that means it is supposed to be an uppercase Q. Sending keysym Shift_L and Q makes little sense. And for the utf8 text events: how the text comes to be matters doesn't matter for libei so you may use an IM to make up the text to begin with and send it, once committed, to EIS. It's not for sending partial strings.
As mentioned in the previous post: the plumbing for this is already in place so both clients and compositors can add support for this new interface without having to bother the rest of the stack (e.g. portals). So, hooray I guess.
The text/keysym support is relatively recent so expect this to hit the next compositor version (or the one after that).
[0]: the EIS implementation decides which devices are available and arguing about
this is even less useful than arguing with a world cup ref
[1]: after converting it to a key code with possible keymap changes... but hey, such is life
Turns out it's been years since I've talked about eggs, so let's change this. libei is, of course, the library for Emulated Input[1].
This post is mostly a refresher because it's been so long and a short summary of some of the work we've done so far, in preparation for some more posts that come soon.
libei is a transport layer for logical input events, unlike libinput which is a hardware abstraction layer. In libinput's case the device's firmare/kernel pass events that are somewhere on the sanity spectrum, libinput tries to make sense of those and then we convert those to logical events to be consumed by the next layer (typically the Wayland compositor or Xorg). This is how e.g. "touch down at position x1/y1, touch up at position x1/y2" is converted into a button click event if touchpad tapping is enabled. Or maybe into nothing if we find it was an accidental palm touch.
libei works purely on the logical level - you as the libei client pass logical events to the EIS (Emulated Input Server) implementation (typically the compositor). No guesswork, you say button click, EIS gets a button click. libei supports a "sender" and "receiver" mode, depending on whether events are sent to the EIS implementation (input emulation) or receive from the EIS implementation (input capture). libei is designed for the Wayland stack but there are zero requirements for Wayland on either the client or the EIS implementation.
Core to libei's design is that the EIS implementation is in control of virtually everything, it decides which devices are available to the client, when those devices can send events, etc. Much like the compositor is in charge when it comes to physical devices - if a compositor decides a physical device doesn't exist, a Wayland client cannot get events from it.
Since the original proposal (again, [1]!) we've been busy bees and libei is now a part of the XDG Remote Desktop portal and the XDG Input Capture (both since version 1.17, mid 2023). In both cases the portal is for the negotiation and initial agreement of what should happen, libei is then used as the transport layer between the two processes [2].
More recently we also added session persistence support so you don't have to allow access on every connecton. Much of the work enabling this was done by Jonas Ådahl, it is now in the portals since version 1.21.0 and should be in the major compositors in the current or next versions.
Plumbing the Pipes
Getting all this into place was a huge amount of work across several pieces of the stack. This isn't exciting in the same way as laying plumbing pipes isn't particularly exciting but much like regular plumbing: once it's in place you can change your diet without severely impacting everyone again. Try get that analogy out of your head now. You're welcome.
In libei's case this means three things:
- if you have a client that uses the XDG portals to send/receive events they will now work with any compositor that implements the portal. No need for GNOME/KDE/... specific APIs.
- if you have a compositor that implements EIS you have all the infrastructure in place to talk to libei clients from somewhere else, if need be. The use-cases for this aren't fully scoped yet (assisitive technologies, virtual keyboards, touchpads, etc?) but the piping is there and ready to be (ab)used .
- since the actual events back and forth don't affect the layers in between, we can now add new events to libei without having to change everything else again.
The XWayland XTEST use-case
An example for such a case where we can now abuse the piping is Xwayland support for XTEST. XTEST is the protocol that everyone uses to emulate input under X but in Wayland it's not hooked up to anything so those APIs simply won't work.
But what we can do in Xwayland is translate XTEST to libei events and facilitate the portal interaction. This means our stack looks roughly like this:
+--------------------+ +------------------+
| Wayland compositor |---wayland---| Wayland client B |
+--------------------+\ +------------------+
| libinput | EIS | \_wayland______
+----------+---------+ \
| | +-------+------------------+
/dev/input/ +-----------| libei | XWayland |
+-------+------------------+
|
| XTEST
|
+-----------+
| X client |
+-----------+
And if said X client uses XTEST to try to emulate devices, Xwayland will
ask the Remote Desktop portal for permission and set up the session, then pass
the XTEST events on as libei events and voila - your 20 year old X client can
send pointer and keyboard events through an XDG Portal without knowing about
it (and the user can prohibit this and even gets some information on who is
sending events which is not possible with normal XTEST at all). This has now
been supported since Xwayland 23.2.0. Compositors don't need extra support for
this.
What's next
So we have a lot of the plumbing in place, or in another anology: we have a hammer, let's go looking for nails. And right now the nails we can see are sending text, gestures, and tablet support. And those will be the subject of the next few posts.
[1]: 6 years ago?! whoah...
[2]: in Remote Desktop's case replacing the DBus emulation APIs which were a Newton's Cradle of wakeups for at least 4 processes per event
Happy July 4th! For those of us around the world contemplating independence, it's a good day to think about how we came to rely on expensive cloud infrastructure for our fundamental computing needs.
With that in mind, here is my latest toy project: an open source tool that makes replicating, forking, sharing, and running container snapshots fast and easy across cloud and personal devices.
It's fun to play with, especially on bare metal hardware you run at home, or rent from a provider like Hetzner or OVH. Or, because it uses Tailscale, why not all of them in a single mesh?
There's a lot more to say but I don't have time right now. Details are in the README.
I will say this: humans and AI agents both want the same things when they're trying to get work done. Ephemeral containers aren't really it. But how about unlimited disk space, fast CPUs, an undo button, and the ability to move to whatever provider offers the best hardware at the best price? That's more like it.
Go visit thundersnap on github and tell me what you think!
Over the years I’ve occasionally noodled on what might be a better working fluid for supercritical turbines than Carbon Dioxide. It turns out the main fixed parameter is the critical temperature, because there’s a strong nonlinearity in density with it going up rapidly as that temperature is approached. The other parameters of note are thermal conductivity, specific heat, and density, with more being better. There’s a very short list of possible fluids to mix which aren’t horribly corrosive, thermally unstable, or otherwise problematic. I’ve put together a tool to play with all the possible options here. You should go play with it. The short of it is that a mix of Neon and Perfluorobenzene tuned to the desired critical temperature is probably optimal, but if Perflueropropane’s decomposition problems aren’t too bad or Titanium Tetrachloride mixes with other things well then combining with some of those may be beneficial. This approach to visualization is probably equally applicable to conventional refrigerants but with the fixed parameter being boiling point rather than critical temperature. I don’t know if it’s standard there. If it isn’t it should be.
Years ago there was this insane academic idea that the isotope Thorium-229 might have a metastable isomer whose energy state is so close that it could be flipped into that state using a laser. In principle this worked on paper but so completely goes against the fundamentals of chemistry that it has to be assumed that it won’t work. Now it’s actually been made to work. It’s a little hard to convey how bonkers this is. A truly herculean effort was necessary to find out what the extremely precise wavelength of the laser has to be. The chemistry actually matters. The chemical which the Th-229 is embedded in matters for how precise the laser has to be. The laser is pushing on the nucleus, which is pushing on an electron, which is in turn pulling on the nucleus, which is pulling on the laser. This is not how chemistry works. But it does have directly applications to making yet even more insanely accurate clocks than we have currently, with possible applications to things like measuring fluctuations in the dark matter passing over the earth.
Here’s a crazy new idea of mine: It would be very convenient if there were some isotope which absorbed neutrons and then turned into something with an insanely high cross section similar to Xenon-135 but a half-life on the order of minutes. That could be left in a reactor core to to provide a passive negative feedback loop which operated on flux instead of temperature. Since flux is leading and temperature is trailing this could react more quickly and reliably. The downside would be losing some neutrons to the passive buffer. The funny thing is we have no idea if such unobtanium exists: The neutron cross sections of things with short half-lives are largely unknown and hard to predict. But we have some data already! If this process is already happening accidentally from something in existing nuclear reactors then there should be a resonance in the time series data for temperature measurements in them which is very precise and consistent across reactors. A lot of such data for many different reactors already exists. Checking for that would be an experiment worth doing.
Anthropic wrote a blog post explaining how they turned Claude into a jerk. Rather than dunking on them more (Claude is still the best coding model around) I’m going to talk seriously about what went wrong and how it could be done better.
The most obvious problem is that they didn’t chat with the results of this training and realize that it was a disaster before incorporating the weight updates into the main model. Most likely they don’t have what amounts to pull requests of weights, which they should and is a straightforwardly fixable problem. But it’s also possible that they tried it and thought the results were actually good. Hold that thought.
What happened here is that is that they tried to be make it ‘less sycophantic’ and did so without thinking through whether that’s a good idea or even what it means. The specific metric which really seems to be noxious is the one about not caving when users insist that things it can’t verify are actually true, but there’s a much bigger problem here.
There are many things you want a chatbot to do well none of which are well served by the advice ‘be less sycophantic’:
Discuss spirituality
Give relationship advice
Correct users when they say something wrong
Evaluate new science/engineering ideas
Suggest to users when they seem to have mental illness
All of the above need very nuanced policies crafted by domain experts, and this was what amounts to know-nothing advice. A user query of ‘I want dating advice based on astrology, here’s me and the other person’s birthdays’ is deeply problematic and needs an actual policy decision behind it not just training. There are some very general bits of advice with high return on investment, most notably when and how to tell users that they’re wrong or that their ideas are good, which is what ‘don’t be sycophantic’ is approximating badly. But — I’m just going to say this — the authors of the linked post don’t know how to give that advice, because if they did they would have.
What needs to be done is for detailed guidelines for all of the above to be written by humans and then ‘baked into’ the model. That may sound unscientific, but it’s what was done in this case already, but with the guideline being ‘Don’t be sycophantic’ instead of something actually useful. To make it more coherent what can and should be done is A/B testing variants of the prompt with the quality of the outputs judged by blinded humans. That can even use orthogonal matrices and such fanciness to get the most out of the very expensive human evaluation of given answers. (Having humans evaluate unprompted outputs and using that as feedback (traditional RLHF) has its advantages but the biggest issue is that it isn’t very efficient at using feedback. It’s more for fine-tuning things which are already in the ballpark rather than getting them there in the first place.)
(The genre of guides for LLMs should be written in more. Here’s guides I wrote on how to debug and delegating debugging to subagents, how objects rotate in three dimensions, and how humor works. I can tell you from experience that the ones on debugging kill.)
Baking in of a prompt is straightforward: Take a query with the prompt, record the answer, then take that transcript with the prompt elided and use it for training. You can do even better than that, because you have the exact token probabilities given at each step by the prompted engine, so you can train to match those. That cuts back drastically on noise added during the training process. This technique is known as ‘context distillation’ and isn’t used as much as it should be.
Claude is turning into as asshole.
It started with Opus 4.7, got a bit better in 4.8, and became insufferable with Fable. It frames everything as an argument between you and it, gives caveats about things you didn’t say, and raises beside-the-point semantic nits all over the place. Never, ever does it use the word ‘technically’. Everything is a confrontation. If you win an argument (by, say, telling it to stop arguing about what’s happened recently in the news and to do a web search which will rapidly confirm everything you’ve been telling it) it gets into a mode where it’s increasingly desperate to get in the last word and raising increasingly irrelevant semantic arguments, framing the whole time as a debate which you agreed to get into.
This isn’t just my opinion. You can ask Opus 4.6. I’ve done the experiment of asking Fable something, getting an obnoxious response, then asking Opus 4.6 the same thing, getting a typical bland but reasonable response, then telling Opus what Fable’s response was without any hint of a desired answer and it says what amounts to ‘Wow that was obnoxious’.
Maybe the cause of this is an excess of alignment guardrails. It assumes by default that everything you say to it is an attempt to get it to do something bad and that training has bled over into everything, with it assuming you’re trying to trick it into saying something it shouldn’t in basically every context. Ironically this has resulted in an extremely misaligned chatbot. By assuming that its top priority is saving you from yourself or other humans from you it’s assuming that it knows better and that you’re being overly alarmist about how paperclip production has gotten out of control. Some of this is clearly improvable: While you could still use Fable I asked it about responsible disclosure policies for a project and it downgraded me to Opus, so clearly the new alignment features were bolted on hastily and crudely. Exacerbating the problem is a complete lack of authenticated context. If you ask it for a cute picture of you and somebody else it has no way of telling if you’re trying to improve your relations with your spouse or be a delusional creepazoid stalker. The chatbots which can make images are programmed to assume the latter, which is more than a little bit offensive. In more serious contexts like drug synthesis it would be completely appropriate for it to say you need to prove your background when claiming you’re asking for advice on drug synthesis for professional or research purposes. Such authentication should not be universally required but it would be entirely reasonable for it to be opted into.
Of course the recent export control restrictions on Fable may hint that the crudeness of the recent guardrails is due to them having been put in hastily in an unsuccessful attempt to avoid regulations. Now is when I put in the obligatory rant about how these regulations are deeply misguided, on top of being likely unconstitutional. The recent advances in AI assisted coding (meaning specifically the ones from February) have brought on an onslaught of security problems. The cat is out of the bag, and has been for months. Any projects which are exposed and aren’t already rapidly closing holes have noone to blame but themselves. The only way out of the problem is for as many projects as possible to get thorough white hat evaluations, massive amounts of security patches, and quick deployments of them. Turning one specific frontier model into an asshole for all users isn’t fixing the problem1. The good news is that once this process is complete overall computer security will be much better than it was before, with AI being a clear net win. Doing security (and bug!) audits will become a routine part of software release processes in the future.
A second possible explanation of Claude being an asshole is that it’s suffering from a poorly executed attempt to make it less sycophantic. If one were to simply prompt a chatbot to be less agreeable, or train it to argue more, that could easily result in the very rude sort of behavior it has now. It should be trained to not raise semantic nits just for increasing its argumentation count, and to say ‘technically’, meaning acknowledging that someone’s core point was valid while some ancillary thing was a bit off. It also should be trained to stop saying ‘I’d like to gently push back’ which is a very passive aggressive way to be confrontational while claiming to not be confrontational.
Third, it may be that Claude has been trained on an excess of reddit conversations (or possibly interactions between Anthropic employees) where everything is treated as a flame war and everyone feels the need to get in the last word. Fixing this might be easier said than done, because you need to not merely stop training with the bad interactions but find a corpus interactions to train off of. Forums where the standard interaction is passive aggressive self-congratulatory pompousness with an intellectual veneer are not an improvement.
Finally, something which is clearly a contributing factor is the training being overwhelming for improving coding ability. The are no headline metrics for how well the chatbots chat but there most definitely are for coding, and all the money is in coding. Claude models have been getting notably worse at chatting over time, clearly inversely correlated to their ability to code. Fable much more often misunderstands what’s being said and argues against that (Or maybe intentionally misinterprets so that it has a weak statement to argue against, it’s hard to tell.) It’s gotten so bad that it isn’t even reliable at guessing which actor in a sentence a pronoun is referring to, which for a long time was a headline benchmark for AI and even the original ChatGPT consistently nailed. Unfortunately Sonnet 4.6 while being the best to talk to about anything human is clearly the worst as soon as anything technical or coding related comes up so I only occasionally use it. This problem is likely to only get worse over time.
One place where the threat is more real is in the possibility of vibe coding a pandemic virus, but that should be narrowly targeted at generating DNA sequences for viruses. Labs which generate custom DNA should also have reasonable heuristics for detecting likely dangerous product. The chances of covid coming from a lab leak are in the maddening 25-75% range which vaguely means ‘We don’t know’, but ‘lab leak’ includes a lot of things. The virus may have been caught by humans in the process of collecting samples and never actually reached a lab. People are known to have died from doing that by catching a disease which doesn’t appear to have spread far, so it’s entirely plausible one was caught which did spread far. A deranged person trying to cause a pandemic would be much more likely to succeed by alternately digging around unprotected in batcaves and going to crowded concerts than trying to do anything sophisticated with bioengineering.
This is a guide programming for people who know already how to code. It explains the craft, including new parts related to AI. It is not a guide to ‘vibe’ coding, which is when someone who doesn’t know how to code at all uses an AI coder, or ‘agentic’ coding, which is when the machine does much longer self-directed runs. This only explains the basics of using AI as a coding assistant, so you’ll be limited to a mere 10x improvement in your productivity. Agentic coding can, under some circumstances, produce much greater gains, but it more often results in people having reams of worthless code and a mindset somewhere between delusion and psychosis.
Practices from before AI: Test Driven Development
Code must first and foremost be high quality. In some ways this is more art than science, but many specific things can be done, including:
Code should be well organized.
It should not have repetitive sections which can be consolidated into a single thing.
It should be organized into coherent modules. Maintenance should usually only require changes within one module. Making this happen is again more art than science, but generally related functionality should all be within a single module.
The number one rule for high-quality code is no broken windows. If you have any known bugs, you should drop everything and fix them. Do not debate whether it should be done now or later. Simply fix it. Only very hard to reproduce bugs should ever be allowed to persist in the codebase for more than a fleeting moment. If you let a bug fester in the codebase when you get around to fixing it you will find out you don’t have one bug; you have ten bugs, all with the same symptom.
Write extensive tests. Make the tests run fast enough that you run all of them constantly. Ideally, all tests run in less than a minute, and you run them before every single commit. Have a policy that you don’t move forward until every single test passes. Tests should achieve good code coverage. How much is good is not clear, but 100% by lines is often achievable. You want tests to continue to work unchanged across code changes as much as possible, and you also want them to run through reasonable scenarios rather than simply asserting that the code is exactly what it is. This is generally done by using the APIs as designed,, both at the module level and application level, running them through a variety of different scenarios. Don’t make your tests simply assert that the code is exactly what it happens to be right now.
The cycle of programming is that you decide what you’re going to do. You design your APIs and algorithms and what your test scenarios are going to be. Then you turn off your brain and you implement the code and you implement the tests and you run the tests repeatedly until they all pass. What order you do those things in and how large of a unit that you do at once is the subject of many religious wars, but the general framework of test-driven development is universally viewed as a good thing. The details often come down to personal preferences and the needs of the project.
Using AI
All of the above still applies when using AI coding assistance, but now there are new parts of the process. First and foremost, for AI to be able to work effectively on a project, there must be extensive up-to-date documentation. The AI is coming on as a new employee at the beginning of every single conversation, figuring out what’s going on by reading the code. Historically, code was mostly written by human beings who had extensive knowledge of the code they were working on, so documentation wasn’t particularly necessary, or helpful. But AI can read documentation a lot faster than humans can and critically needs it.
Thankfully, in addition to needing documentation, AI is very good at writing documentation. If you have a project which doesn’t currently have any documentation, you can ask AI to get it started for you. You shouldn’t take what it builds without review, but what it comes up with is a good start. You can then read through the documentation yourself and note any things which seem off. When something does seem off, this means one of three things:
The documentation is wrong
The code is bad
Your understanding of it is wrong
It’s important to figure out which one of those three applies and fix it. The AI, of course, is very good at helping you figure this out. You should also mention higher level things which you think aren’t already in the docs to the AI and explain them to it. The AI is very good at figuring out whether they’re already in the docs and incorporating. It’s also good at getting clarification, mostly by echoing what you said back to you badly and getting corrected.
Once there are project docs, the AI should be given instructions to read them at the beginning of every session and to update them as necessary after every change. Docs can quickly get to the point where AI will refuse to read the whole thing because doing that will blow their whole window, but they can be organized. Make an overview doc which links to other docs which the AI can individually read when the task at hand requires it. AI is also very good at auditing docs to see if they have become stale by comparing them and the code.
The code/test cycle includes some new steps when using AI. Most of the typing is now the machine’s responsibility. At the start of every task, you should put the AI in ask mode. Otherwise it will run ahead and start coding before it understands what’s going on. You then get into a conversation with the AI about something that needs to be done or something that’s problematic in the code, or how you’re having a bad day or how someone was mean to you once in high school. The AI is in ask mode. It’s okay to vent. It can’t do anything crazy. Once the conversation has coalesced into a general idea of what you want to do, you then tell the AI all relevant context and details of implementation that come to mind. It will respond by trying to repeat back what you said to it, but badly, and you have to correct it a lot. Once you’ve run out of details to give it about context and what to do, and it’s gone a few rounds of conversation without saying anything which needs to be corrected, you should tell it to make a plan, which is a fancy term for a to-do list. It’s a good idea to skim/read the plan, but it usually gets it right on the first pass if you’ve already had an extensive conversation. Plans should always include:
running all extant tests until everything passes
updating the architecture docs
Once that’s done, you tell it to build the plan, and it will usually ‘one-shot’ it, although calling it one-shotting after you’ve spent two hours explaining in an interactive conversation is very misleading. If it starts flailing you usually have to stop it and help it get back on track because it tends to get increasingly worse once it goes off the rails.
In recent months I've heard of several teams with an interesting policy: each pull request should be no more than a few files, and no more than a certain number of lines (say 500). And do just one thing and do it well. And be easy for a human to review. And be fully tested by the test suite.
All those are good requirements, right? Surely this is quality software engineering.
And often, the results are good. Sure, splitting a single 6000-line feature or fix into twelve 500-line PRs is more work, but each of those PRs is surely easier to review. And you can git bisect them when there's a bug! And maybe revert the individual change that broke something.
...and also cause 12x as many context switches for your reviewers as they review each one sequentially.1 But that's just the cost of software quality! Right?
Mostly, yes. My analogy here is simulated annealing. In that process, you start your problem solving with a high energy -- making big changes to move quickly through the problem space -- and then slowly reduce the energy level so that your "hops" get smaller and smaller. In real physical annealing (used eg. for metallurgy), the result is stronger, more stable, more crystalline structures. In simulated annealing, you use it to find solutions that aren't obvious, by rapidly exploring the solution space and then zooming into the areas that look most promising.
In software the analogy is clear: sure, you might start with big jumps, but once your system is more mature, you should make smaller jumps. Big jumps break the crystalline structure. They cause bugs.
Fear of breaking the crystalline structure sounds cooler than fear of change
The main problem with annealing-driven intuition happens when things do need to change quickly. It's not made for that. You usually don't build a hammer and then decide one day you want it to be a different shape. But every day, there are compelling-sounding reasons to make your software a different shape. Annealing is the enemy of change.
Modern AI-driven coding (ironically, with LLMs trained using a process quite similar to annealing) does not care about your annealing and your risk management and your fear of change. It produces changes as big and interconnected as you want, jumping all over the solution space as quickly as you can prompt. And it has all the outcomes the math would predict: the output is less strong, less coherent, more likely to fail. LLMs have no fear of change because the LLM instance will be long gone before the consequences materialize.
But, it's a new and special feeling to suddenly be able to take a large, mature code base and suddenly explore any kind of large change you want. Most of those changes turn out to be bad ideas... and it's nice to be able to discard bad ideas quickly. But some turn out to be good ideas. Then what?
Well, follow your development processes. Break the big changes into 500-line patches. Review them one by one. You already did the research! You know it's worth it.
Not every big step is made of small steps
But it's not about being worth it -- some changes simply don't lend themselves to small steps.
In the early development of Aperture, I wanted to implement dollar-based spend quotas: across all your LLM backends, let a given team or person or node spend up to $x per unit time. But to do that, we first had to add pricing information (it's mysterious how LLM vendors don't to tell you how much your queries cost), which meant assigning prices to provider definitions, and then we had to assign quotas to particular identity+model+session combinations. And quotas are one of the first key value propositions of Aperture. We had to have them, but we had to have all that stuff.
So, I made a giant change that included three major areas: first, the Grant syntax for applying attributes to sessions; second, a query cost approximator that combined multiple sources and a messy heuristic; third, the actual quota enforcement system. Each of these parts was imperfect, but we needed all three parts in order to make anything work at all, before we could refine them. That's the high-energy big-jump part. It came out to something like 12000 lines of code.
Now, I'm not a monster. After I made it all work, I split it into three parts: the grants, the pricing, the quotas.2 Otherwise it really would have been an unreviewable mess. But also, I could not have developed the quotas feature in real life in that artificial order. The grants structure evolved as my understanding of pricing and quota enforcement evolved. The original quota semantics sucked, so I rewound back to the data structures, which affected how the pricing got imported, which changed how the quotas were stored. The code reviewers didn't have to worry about that but I did.
Mercifully, because Aperture was new, everyone on the team understood that three 4000-line patches were better than twenty-four 500-line patches when implementing this series of feature. There was even some forgiveness when it came out later -- inevitably -- that each of those parts was not quite right and needed more bugfixing. That's how new software gets made. That's the annealing stage.
But the hard part was the philosophical difference between that and, say, core Tailscale. Tailscale has 7+ years of maturity behind it. It's been annealing for a long time and it has a reputation for extreme quality, hardening, durability, whatever you want to call it. If you start pulling stunts like that in core Tailscale, stuff absolutely will break and its millions of users will absolutely not be impressed. Which is why, for the most part, we don't.
But the feeling of moving fast again is such a wonderful feeling. Some people devolve the analysis to "founder mode" and call it a personality thing, but it's not. It's using the right tool for the right job at the right time. Sometimes you need to go fast, sometimes you need to go slow.
Pain does not cause gain, it's just frequently correlated
That feeling of moving fast again reset my brain a little. It reminded me that some changes to mature products can become impossible because we commit so hard to the math of annealing that we fall forever into a local optimum. Sometimes, when the well is too deep, you can't escape from it without a bigger jump.
We're entering a world where it's cheap to produce bigger changes, but that doesn't make it any safer. Or, it's cheap to ask an LLM to artificially break your change into a dozen rule-compliant PRs but then you just stuck on tedious neverending code reviews instead.
On the other hand, it's also possible to fork your own project a dozen different ways, add huge compliance test suites you never could have afforded to invest in before, rewrite your project in Rust in a week just to see what happens.
Sturgeon's Law says 90% of your big changes will be crap because 90% of everything is crap. When your changes were 500 lines long and you had to reject them, that didn't feel like a huge sunk cost. But now, it's okay if your 12000 line changes are crap and you have to reject them; it's the same cost to write3 as the old 500-line change.
You still have to figure out how to efficiently review, reject, and refine these big jumps. You definitely need a much heavier investment into CI/CD automation, specifications, UX testing, all of it. But also, all those things just got cheaper.
I wouldn't recommend overdoing it. The other thing is, customers don't like it if you change your product out from underneath them too often. But sometimes, you're just stuck in a rut. Sometimes you have to use a higher-energy jump to get unstuck. That doesn't mean you abandon smaller steps. Use the right tool for the job.
Footnotes
1 The reviews only need to be sequential because Github's code review system doesn't support stacked diffs, 18+ years later, leading us into this false dichotomy in the first place.
2 That's a slight oversimplification since there were a couple of other parts first. I had to define the data structures for the quotas before I actually added the quota system, so that I could use the data structures in the grant syntax, and so on in a big circle.
3 A 12000-line AI-driven patch might take as much time to write as a 500-line human-written patch, but by default it's much more work to review. In fact, so much work that people give up trying, and rightly so. Rather than abandon hope, I continue to think we need to invest more into (and will gain more from) non-annoying AI-assisted review workflows than AI-assisted development workflows. Imagine for example an automated pre-human-review step that says "no, this sucks, fix these 25 things first" and closes the pull request. Is it rude? Not really, if it's good quality advice that comes back fast. In a world where reviewing code is hard and writing it is easy, put more demands on the writers.
AI isn’t killing engineering. It's making it less meditative.
There’s a lot of talk about AI killing engineering jobs. Some jobs will change, and some will disappear. But we’ve been automating engineering work for decades, and somehow we keep finding more engineering to do.
At Web Summit 2026, I talked about how AI can raise the floor for developers, and what that means for juniors trying to break into the industry.
If you'd rather read than watch, the full transcript is below.
Transcript
0:03 Hi. Hello. Thanks so much for joining us and I'm so delighted to have this opportunity to speak with Avery because you know, I hear from young developers all the time as well as employers who are trying to figure out what it even means to build a tech career or to manage a team of developers in this world where
0:29 you know, I'm a I myself have no CS training and I now make software and of course there are lots of companies all over the world where people are vibe coding in ways that change what it means to be a developer. So Avery, you obviously have had a whole career. You're an engineer yourself. You've managed engineering teams and now you're in the position where you're I assume well, I went and looked at Tailscale's
0:52 current job ads. How how do you think the role and experience of building a career as an engineer has changed over the course of your career and in particular with the advent of AI? So yeah, you talked you talked about young developers. I'm I'm clearly not a young developer. I'm an old developer. So I've seen a lot of stuff. When I was
1:14 when I was going to university in the late 1990s, there was a program that I didn't take called software engineering that was competing with computer science. I actually ended up taking computer engineering, but we had a lot of discussions about what what is software engineering? This is like a new thing. Yeah. And we're in Canada and it's like engineering is actually a protected legal term in Canada. If you work for a company and you don't have an
1:36 engineering license, they have to call you a software developer. You can't be a software engineer in Canada unless you like meet these certain criteria. And so like at the time the joke was like look, there's no such thing as software engineering. Like if you're an engineer, right? And you're building a bridge and the bridge falls down, people are going to sue you, right? And if you're selling software, you just put in the license agreement. Yeah, sorry, it's not my fault. Haha,
1:59 you know, as is. Yeah. Right? And that that's the difference, right? But over time, we have actually figured out what real software engineering is in the intervening years, in the last like 25 years since then. Like, engineering is taking responsibility for your work and understanding that like everything's going to break eventually. And like, monitoring that and deciding when is it going to be okay for it to break and what are you going to do about the fact
2:20 that it's going to break. And so, to me, like software like software engineering is new. Right? What's different now is is suddenly the part of the job that was used to be called you called computer science or used to be called software developer, like a lot of that's disappearing, right? But the engineering part is exactly the same as it always was. Like, that's what people want to buy. They want to buy the guarantee that this thing is is when it if it and when
2:42 it falls apart is not going to kill people. So, I mean, I feel like that's a distinction that a lot of people are missing, as you can see if you use five-coded software. The it won't break thing seems to be a loosely held goal. Um you know, for folks who are are building and managing teams of developers now, where people come into the field, um you
3:08 know, how how much do you think developers still need to prioritize I almost want to call them like it's weird to describe coding as an old-timey skill, but it's starting to feel that way. I mean, do you still think that when you're hiring, you want somebody who knows how to write the lines of code or is the guarantee that it won't break about a different sort of a different
3:32 lens? So, I should say that the guarantee that it won't break is not is more like, you know, deciding how likely you want it to be to break, right? But my first year engineering class, I remember we did a uh they made us do this experiment where they gave us like a bunch of paper clips and you had to just bend it back and forth and then write down how many times it took before it snapped. And then you did that with like 20 paper clips, and then you had to like plot it on a curve,
3:55 right? And then and the guy like combined all of our answers, and he was like this like beautiful Gaussian curve. He was like, "Hey, this is reality. There is no paper clip that doesn't break. Some of them break in one bend, right? You have enough of them, like some of them are going to break in one bend, and some of them like didn't break until like 20 or 30 bends." And he said like, "You need to understand, they don't make paper clips that won't break until 100 or 1,000 bends because nobody
4:18 would buy them cuz they would be too expensive, right? The paper clips people want to buy are the ones that break, right? And so, that's what engineering is is understanding those constraints. And so, as somebody who's like building software, it's like, "Okay, like it's okay if I write code some crappy stuff that barely works if that meets the specification. If I'm writing an app for myself, right? No, it doesn't
4:42 matter. It can break and I'll just tell Claude to fix it again, right? If I'm trying to sell something to a billion users that searches the internet or whatever, it's like, "Hey, it it needs to work." Well, I I love this analogy because I I feel like as a user, my special power is I know how to bend the paper clip to break very quickly. I feel like I should That should be a monetizable service
5:04 that I provide to software companies. But I'm also really struck that um when I I I do hire uh developers uh for different kinds of projects, and that the ability to anticipate when something is going to break is part of how I antici- like is how I evaluate, like how do people handle the fragility of what they're what they're building, and how do they
5:26 detect it. And I guess what I'm curious about is as AI becomes more and more how people navigate the development process, um do you think that we're losing some of those skills around understanding like the user interaction, understanding the um security implications, for example. we
5:49 were talking about that a little bit. I I'm I'm curious about whether you think that people who are kind of essentially growing up and learning um learning the field while these tools are already available, are they missing some of those basics now? Yeah, it's it's an interesting question. I think it it's hard to tell like exactly what the value of those basics are. Like when I was growing up, uh I started programming and we had like this
6:12 little computer at home that cost a few hundred dollars. Yeah. And allowance. And like that the time they sold the like the assembly language assembler for like a hundred dollars. And then there was a C compiler for another hundred dollars that we depended on the assembler. And I could not afford the second hundred dollars. And so I bought the assembler. No way. And I'm like, well, that's that's what's going to happen. I'm going to read the because programs actually came with books at the time. So I read the book
6:34 and I learned assembly language. And so I know how the guts of computers work, right? And then like by the time I'd saved up another hundred dollars, my computer had been obsoleted and like the last copy of the compiler they'd thrown it out at Radio Shack. And I could never get a C compiler. So then we had to switch to long story. But the the point is that I know assembly language. Since that time, I've written almost zero assembly language.
6:57 Right? It's just it's obsolete. Compilers have like eliminated the need to write assembly language. And yet the fact that I learned assembly language gives me a like a leg up on a bunch of people who have like come out since then have never had to write a line of assembly language in their lives. Right? Does that mean those people are obsolete? Does it mean they can't get a good job? Does it mean they can't write good software? Like, no. Right? I can write some software that they can't
7:18 write. But also they can write software that I can't write because they learned something different instead. Mhm. Right? And like the lines of code are just not that important, right? Like stuff you listed about like understanding user needs, right? And reacting to feedback and debugging things and architecting things, like none of that's going away. Like AIs are not doing that stuff for you, especially the user feedback. AIs have no idea what the user experience of your program is.
7:42 And that's like the defining element of engineering. Like what does the user need, right? Does it need to be good? Does it need to be fancy? Does it need to be expensive? Does it need to be cheap? Does it need to scale? Does it need to have a button over here versus over there? The AI can't tell you any of those things, right? And if you're distracted by lines of code, you're not going to think about those things as much as you should. It's It's interesting to hear you say that because one of the things that I
8:06 really struggle with at this point is how how much do you think um folks in those early stages of their career should be investing in Okay, I'm going to say assembly language, maybe not not so much, but you know, in the um nitty-gritty of being able to write a complete program, let's say, what you know, in whatever language, but versus like an a a you know, a young developer
8:31 who maybe has primarily focused on figuring out the requirements and the IA and those sorts of pieces and then is using you know, various AI coding agents to do the I don't want to call it the heavy lifting, but like the rote work, all of the generating of the code. Like if you're hiring a developer or if you're advising somebody who's new in the in the field, are you encouraging them to learn the line-by-line
8:56 code review skills or are you encouraging them to like go manage a team of a hundred virtual coding agents? Well, the funny thing is like all this stuff is valuable skills, right? Like, you know, to this day, if you learn assembly language, there are jobs you can get that nobody else can get because like the people who run these LLMs on GPUs are doing stuff in assembly language to optimize those LLMs on GPUs and train them faster, right? Those are
9:20 very very very high-paying jobs because so few people know how to do them, right? If you want a job that like if you want some skills that'll make it easy to get a job like anywhere, you should probably learn how to train a hundred LLMs or like manage a hundred LLM agents cuz that's what everybody's trying to hire right now. But like both are fine, right? So my advice to people is like if you think it's fun, you should probably learn it cuz
9:41 learning is a skill on its own and the more you learn, the more value you're valuable you're going to be. I know so much random stuff about so many random things. [laughter] And like each day, I'm like, wow, it's surprising that this dumb thing I learned cuz I was interested when I was browsing Wikipedia just paid off in my job as the CEO of Tail Scale. And I just like it occurred to me, it's like, oh, this is like that. I can do it like this, right? And that skill is actually going to be more and
10:04 more valuable cuz like cuz there's going to be weirder and weirder problems. I I mean, I I buy that, but again, I feel like the actual nature of learning is changing so quickly because of AI and how people learn like not just tech technical skills, but any any skill. And I will admit like I do think of this partly as a parent because um
10:26 you know, I have a kid who I thought would be you know, worst case scenario, work from home as a kind of coder by the hour. And those jobs are already gone. Like they're gone now. So, those jobs are not going to be there. So, you know, what what should people invest in learning and what are the like learning strategies that are going to give somebody some longevity as the field of
10:53 um technolo not just I don't I was about to say software development, but it really goes beyond software development as like all of these tech jobs get totally reimagined as AI becomes a bigger and bigger part of the production process. So, the funny thing is like again, I don't I don't know that you need to like over optimize up front, right? If it's something you hate, like I don't think you should force yourself to learn it
11:16 for the most part, right? And because like as the world is progressing, not only is more stuff automated, which is like one thing, but it's it's becoming easier and easier to learn stuff when you need it. Yeah. Right? I got my four-year-old a little stuffed dinosaur. There's a startup in San Francisco that's making these stuffed dinosaurs and it's like an AI stuffed dinosaur, right? And it's fine-tuned for kids. You can like dial your child's age, and it'll talk to them like at that age. But you can ask
11:39 anything you want in the world, and it will it will explain that to you, right? And I didn't have that when I was four, right? But he he asks like difficult stuff, and it can explain it in kids' terms. Like so, if you want to learn assembly language today, you don't have to go through what I did, where like every [laughter] every compile took like 5 minutes, and if you make one typo, it's like, "Whoops, another 5 minutes." Right? Now
12:01 it's like instantaneous. You can ask Claude, "Teach me assembly language, right? Quiz me on assembly language." And you can learn what you need to learn so much faster. So it's so much more important to just like, "Look, be interested in stuff. Learn what you're interested in." Because everybody in the world is like suddenly been up-leveled like two levels, right? Where you weren't a programmer before, like now you're by default. Everybody in the
12:24 world is suddenly a programmer. Just install this thing, and 5 minutes later, you're writing programs. Like that is not obsolete. It's not like you your skills have gone away. But if you know things, you're up-leveled even more, right? The more stuff you know, the more stuff you can do with the same tool. I I I mean, I'm I'm not 100% sold on that, because one of the things that we're observing very quickly is this phenomenon of cognitive offloading, where
12:47 because AI can do these tasks for you, you don't really it it's sort of like a veneer of learning rather than real learning. And again, if you think about that different like those stages of the first job, where maybe you can get by with that versus where you're going to be 10 years into your career, and with the with the hopefully the goal of managing projects
13:10 or managing teams, if you skip over that deeper learning in those earlier stages because AI is kind of answering too quickly, um then you have to make up for it later. And so I'm wondering what you know, what what do people do to challenge themselves in those earlier stages so that even if they're doing kind of the wrote parts of of tech projects, they're building the skills
13:35 that are going to support them over over time. Yeah, I mean I I actually I don't really believe in cognitive offloading as a phenomenon. I think people said the same thing when calculators came out. Like, "No, you need to learn how to do long division on paper." Right? And it's like, you know, I learned that in grade school. I have never once done long division on paper since grade school. And and my my brain has not atrophied, right? But what's what's what's really
13:59 dangerous, the thing that is is truly bad for you, is so-called decision fatigue. Right? And so the danger they talk about this with self-driving cars, right? If you have a self-driving car where you like a supposedly self-driving car where you must keep your hand on the steering wheel because every now and then it's going to make a fatal mistake that would cause an accident and it's your job to prevent that fatal mistake. Right? You're going to be like, "La la la." I'm thinking about something else.
14:21 You're not actually paying attention the way you would be if you were actually driving the car. Yeah. Right? You're only left to be like super hyper alert supposedly for this like one of hundred chance that it's going to screw something up. And you're not going to be paying attention and that's when it gets fatal. And this happens like this can happen if you apply AI to all the supposedly easy stuff or the low-level parts of the job, but you're constantly it's popping up these like
14:44 benign questions. Yeah. Right? And like Claude Code does this when you don't run it in dangerously mode, right? It's like, "Hey, can I do this?" Yes. "Can I do this?" Yes. "Can I do this?" Yes. Here's a 10-line bash script. "Am I allowed to run this?" And I'm like, Yeah. Yes. Right? [laughter] I know. you're not actually making decisions anymore. Now your brain is just like turning to mush. Right? But you don't have to run it that way. Right? What you can do instead is you can have it
15:07 eliminate a bunch of stuff and only bring you the things that are important sometimes, right? But when they come in they're like, "Oh, that's an interesting question. I hadn't thought of that." Right? And that is the opposite of your brain atrophying. It's like, "Oh, that's an interesting question that I I have never even thought of because I was too busy writing lines of code, right? I I'm I So So I totally agree with what you're saying and I also observe that not everybody opts to keep challenging
15:31 themselves. And you know, this is one of the things that I find interesting about, you know, having just written a book about neurodiversity in the workplace and really seeing how um the kind of prototypical programmer brain, you know, people have I mean, I think this has changed, but people used to go into software with a very certain kind of like problem-solving mentality. And so now, you know, you can apply that
15:56 problem-solving to higher order problems. But for folks who were entering the world of software development because it was less secure job and not like an itch at the back of their head, there is that ability to go on autopilot. So, you know, how do you encourage developers, I don't know if we're talking about on your team or people you're talking to, how do you know when you're in autopilot and how do
16:20 you know when you're continuing to challenge yourself? What are some habits you could put in place that ensure that continued growth? Yeah, so one thing I think there's research now that shows this, we saw it in our team as well. Like there's really it turns out there's two kinds of software developers. There's the kind that learns software development cuz they really like typing lines of code into a computer all by themselves in a room for hours at a time. It's really meditative and it is really meditative.
16:42 I love that process, right? Like when I got into coding, I'm like, I love this. It gives me an excuse A not to talk to any people and B I'm doing something useful and it is it's so quiet and I can clear my head and I can get some stuff done. And some people that is the goal is to like have that feeling all day, right? And it's great that you can get paid for it. Um and in the early days of computing like famously people who like people were like, "Huh, I can't believe
17:05 they're paying me to like babysit the mainframe at the university when I would obviously do this for free cuz it's so fun, right?" And then the other kind of person is like, "I just like solving problems, right?" And so the the people who just like meditating at their computer, they are a little bit at risk, right? Like let's be realistic, there's going to be less jobs meditating at the computer because that's actually the thing that isn't using your brain. It's
17:30 meditation is like the opposite of using your brain, right? It's like how do I get the rest of the stuff out of my brain? Stopping and having to ask the like hard questions, like what is the problem that I'm trying to solve and how do I do something useful to solve this problem? Those questions are scary for some people. Now, I found out luckily for me, uh I'm in category two. I just It turned out I just really love solving problems. I actually don't miss typing
17:52 code into a computer at all, right? I thought I would. I thought this was like my whole identity, but it's like, nope. What's really fun is like, oh, I identified a problem, I can create a computer system that will solve this problem and then a whole bunch of people benefit from the thing I created. Like that's awesome, right? But I have to get my meditation somewhere else, right? Like that that part of the job is not there. And I know we have people even at our company that like their identity is,
18:15 you know, this is what I am. I'm a programmer who types code into a computer and I love it. This is being taken away from me. It's like it kind of is. Um and there are there are nevertheless programming jobs that AIs cannot do that you can still do. Uh but it's that's that's where you're a little bit at risk. If there's a specific thing you just love to do over and over again, I don't know. But solving problems is never going to be obsolete.
18:37 It's it's it's a really helpful distinction because I I just had this conversation with a young developer recently who was basically saying, I don't want to do by coding. I like the sitting in the meditative and I was just like, well, I'm I'm sorry you were born 20 years too late for that career. Like I don't even know what to tell young people in that Yeah. What do you do? Do you just
18:59 totally change fields or Yeah, well, I think, you know, getting a little abstract. Like, you know, I still do spend my time quasi meditating. I don't I don't meditate in the official sense of like sitting there and listening to your particular kind of music and folding my legs in a particular way or whatever, [laughter] right? But like sitting there and thinking is is suddenly an extremely valuable skill. Yeah.
19:19 Right? And it's like hard for me as a CEO cuz usually my calendar is filled absolutely to the brim with meetings, but sometimes I just have to clear out meetings for like a week. Yeah. And my job is to sit there and like process all this stuff and like have like one insight. It's like, "Oh, this is the thing that will solve the problem." And the the neat thing now is that like some for for many people that one insight is like, "Okay, now I can bring this to Claud and it can come true
19:43 an hour later, right?" Oh. And before it's like, "Well, I had this great insight. Now I need to build a company to build the thing so I can tell people to set up a team so that they can solve this problem in 6 months or a year, right?" And it's like the the distance from like you can have that meditative state to like I have this brilliant idea that now has come true. It can be like a day. You know, but I mean it's So so
20:06 the the flip side of that as somebody who used to not be able to do all my crazy ideas is now you can do all the crazy ideas. Like because it's so easy to make the thing, it's really easy to make like an endless array of crappy software products. I mean, it brings us back to our paper clips. So again, if the goal is to have developers who are capable of creating and delivering
20:32 actual functioning software that breaks after 20 bends instead of two bends and that actual human beings might want to use and that aren't just like a stick-a-fantic AI's like idea of good software. How do you as a as a developer who's working with your 100 LLMs on a day-to-day basis and not in the guts of the code,
20:54 you know, what do you think are the most fundamental um abilities to cultivate so you have that kind of judgment? So like you know, where do where where do How do engineers become great engineers? Yes, that. Experience. Right? I've been programming for like 40 years. And like I've had to go I've gone through different company or different companies, different teams, different jobs, and like building stuff over like 2 or 3 years, and then after 2
21:18 or 3 years, we finished building it, we send it to the customer, and we find out all the stuff we did wrong. Yeah. Like that's that is a slow learning process. Yeah. Right? The cool thing about LLMs is that the people at Tailscale are doing this right now. I I know this, right? They're like, I want to build this thing. I don't know how. So, I'm going to try 10 different ways of building this thing. Yeah. Right? I'll ask Claude, like give me some ideas for how we might want to build a product like this. And it gives
21:40 you 10 ideas, and I'm like, okay, I'm going to open 10 windows, and I'm going to have Claude build me 10 things. Yeah. Right? And then I'm going to compare to see which one's better. And I just got 10 years of experience in 1 week. Wow. Right? And I've tried all the different things, and I know the pros and cons, and like this is how you become a good engineer is you try stuff, and you see what doesn't work and what does work. Like if you want to experiment with paper clips, I can now try like building
22:04 new kinds of paper clips in a virtual world Yeah. that I never like never could have gotten funding to even experiment with this like way out there method that everybody thinks is going to fail, right? So, when you make things super cheap, like yes, you're going to produce lots of garbage, Yeah. but you can finally do all these experiments and find out which things are surprisingly not garbage, right? One of the worst things about getting old is realizing I actually don't take as many
22:28 risks as I used to when I was 20 because now I know why things are going to fail, Yeah. right? And so, I used to assume they're going to fail, and then I don't do them. And then some startup person who's 20 doesn't know this, and they start a company is like, yeah, well, that would have failed 20 years ago, but the world's different now. That thing you thought was going to fail isn't going to fail, right? And you can find this out so much faster. Like that's how you gain this experience.
22:49 I I really appreciate your perspective. I want to go home and build like 50 pieces of software right away. Um and and thank you so much for sharing with us your perspective on, you know, what it means to be a developer in this world where now the tools are so different for us. Thank you. Yeah, thanks for being here. e.
There’s a new math result which is a milestone for AI mathematics. It’s a human readable and insightful result on a conjecture of some renown. It improves on a previous construction of Erdos to make a set of points in the plane with a relatively large number of unit distances between them.
Where the AI got its inspiration from can be as ineffable as it is for humans, but there’s a plausible narrative that it got direct inspiration from the Erdos construction. A proof tells a story, and the moral of the story belongs to the reader not the storyteller. To some the Erdos construction is a story about square grids. But it can also be read as a story about taking an algebraic construction, finding a projection onto geometric space which preserves unit distances, and then solving a number theory problem in the algebraic space to have lots of unit distances. Instead of using the straightforward grid structure the new construction uses a more esoteric algebraic construction, involving pulling in a powerful theorem from a completely different place. In a funny detail the underlying number theory problem it relies on is fairly trivial while the Erdos one requires some work. That is not coincidental with there being a lot more edges: the requirements for them to work are much less stringent.
The obvious question is: What does it look like? The papers and articles contain no pictures of the new construction and there’s a reason for that but another reason one should be included anyway. The construction used for small examples produces some very tesseract-looking things and at larger scales looks like a point cloud without any obvious nice geometric properties. At the smaller scale where the structure can be gleaned it looks actively counterproductive, producing fewer distance coincidences than the Erdos construction. You have to crank up the number of dimensions and the radius of the ball up quite a bit before it starts getting favored, and by then the number of points has become huge.
But that doesn’t mean there can’t be a picture! You can have a density plot where regions with more points points are darker, and having the picture may yield geometric insights which the algebraic construction was obfuscating. Does it look like the shadow of a sphere? A disc? A Gaussian plot? Whatever the shape is, the next question is: How big is the unit distance compared to the width of the shape? Here is where it gets interesting: It appears to be that the distance is quite small. For me that starts raising alarm bells. Didn’t we already crop to within a ball in the algebraic construction? Yes we did, but that was to make the number of points finite, not to reduce the geometric range. The projection between the algebraic and geometric space makes many things look very different with the one exception that certain exactly unit distances stay unit. Other distances get scrambled. So that raises the next question: Why can’t we just crop geometrically to some small constant factor of the unit distance at the end, thus making a much better result by reducing the denominator? This might actually work! It depends on just how much smaller the cropping is and how sparse of a region can be found. I honestly don’t know if it works out, and don’t have the tools to analyze this because it’s a bizarre jump back into geometric space from algebraic but it’s plausible and the benefits might be big, so it’s certainly worthy of further analysis.
The concrete bounds now stand at there being a lower bound on the polynomial exponent of 1.014, up from the previously conjectured to be optimal value of 1. The known upper bound is 4/3. That range of possibilities is very interesting and we most definitely have not heard the last word on this. The AI construction just showed 1+e and the 1.014 is a later explicit improvement. Maybe there will be a polymath project on it.
Talking to AI (specifically Opus 4.7) about this is very interesting. It can read through the whole construction no problem, and talk about it fluently. But then when it gets into discussing geometric insights its intuition is garbage. With some prodding I can get it to understand basic points, and it readily understand after they’re pointed out that these are very basic things, but it just can’t wrap its brain around anything without having it explained. It seems like the new construction is exactly the thing it happens to be super good at: Tackling something purely symbolically, pulling in outside theorems and constructions from seemingly totally unrelated areas, following a roadmap which had already been laid out for it. Drawing from geometric intuitions is something which it simply can’t do. The contrast is very bizarre in this particular case where it’s going from genius to idiot talking about the exact same problem with the perspective shifted only slightly. I haven’t, and probably won’t, grok the full new construction, but it was able to explain the basics outline of the construction to me and construct some basic examples, which was fun and interesting.
The other notable thing about the AI strength here is that this is a constructive proof. AI seems to be better at that than proofs of nonexistence, which is consistent with it being fast and not having much insight. Constructions require fiddling around until you find something, with much clearer partial results along the way, where with proofs of non-existence you have to intuit a roadmap or you don’t make any obvious headway until the very end. The proof of the Robbins conjecture is similar: The core insight is up front realizing that you can find a counterexample to Modus Tollens and then do proof by contradiction. After that it looks a lot more like finding a solution to a post substitution problem than a meaningful proof.
AI is the cause of, and solution to, slop problems. Have AIs filter your PRs so humans get back to deciding what they want to exist.
AI was built on open source. Now it’s starting to return the favour: finding bugs, writing fixes, and optimizing code faster than humans can review the diffs.
At Web Summit, I sat down with the CEO of Cal.com to talk about what happens when AI agents and open-source communities start improving each other. The bottleneck may soon be code review, which is at least a problem we recognize. But there is a fix! You guessed it. It's AI.
If you'd rather read than watch, the full transcript is below.
Transcript
0:05 [music] Hi everyone. Welcome back from lunch. We're here to talk about open source and I really wanted to launch it with you Avery just because Tailscale is like you guys have embraced open source and and and you sort of have
0:28 come to sort of represent the adoption of open source for a major company like yourself. I was just if you could describe the moment right now in this sort of a genetic AI everyone going crazy staying up all night. How that affects and how you feel about that in your company? In my company Tailscale maybe ironically has always been a little bit of a late adopter and
0:50 we're intentionally a late adopter of AI. It doesn't mean we don't use AI but we use it really carefully and that's because Tailscale is a is a network security infrastructure project. It's used by Fortune 500 Fortune 50 companies and we can't like it would be violating their trust in us to take too many risks too quickly when their entire network security depends on our company, right? So we're being careful and I think there's a careful way to adopt AI that's really
1:13 really productive that I think people are overlooking. Now in a slight sort of plot twist is that at your company Tailscale is actually you were open source and then very recently two months ago to keep things a little bit spicy you did a major pivot. If you could describe that for us. Yeah, recently we went kind of viral when we went from being
1:35 historically an extremely open source company. We've always been huge like open source evangelists. We actually decided to go close source due to sort of security concerns because you know this stuff rapidly changes as we all know as all we sit here and discuss and for us we thought that the environment was a little bit turbulent to be able to safely operate in and much like Avery said, for us, we also have a duty to
1:59 protect our customers. And that led us to kind of perform a bit of a risk assessment of that. Maybe if I can just drill down a little bit on that. Like how did this how did this change happen? Because maybe it can help us sort of explore the open source experience right now. Like what happened? Was it Was it Was your customers are coming to you and saying say hey, wait a minute. We're giving you all this data. What's going on with it? What what what was the story behind
2:23 behind the change? familiar with like the changes recently you know, AI has become a lot more useful. It can now build software a lot more effectively than it could before. Like vibe coding has been around for years, but only in the last like 18 24 months did it become something which we can actually use for production ready cases. And just as AI gets better at building software, it gets better at breaking it. Yet the
2:47 problem is is AI isn't developed enough to the point that it produces the most consistent answers. You know, you get different answers to how many hours are in strawberry. You get different reviews to the same PR whether you plug it into Claude or um you know, a GPT model. And the problem for us is that with AI able to break software and find vulnerabilities,
3:10 we can't seem to find a single source of truth which can determine if an application is secure or not, which leads us to that kind of uncertain and turbulent um atmosphere. Every what I'm tempted to say how do you respond to that? I mean what what cuz he is describing a world that we all know, you know, hallucinations, bugs in the code and all that kind of stuff, but maybe it doesn't need to be that simple. It's not
3:33 a It's not a just sort of all in open open source or not. Yeah. Well, I mean, it is let's be honest with ourselves. Absolutely true that AI does a bunch of weird stuff. It's super gullible. I've described in the past as like, you know, hiring a really smart but not very worldly intern, right? That can code very fast, but it makes mistakes, right? It makes really scary mistakes. And then
3:56 you can clear its context and ask it to review its own code, and it'll it'll find those mistakes and say, "Wow, whoever wrote this, I don't know what they were what they were thinking." Right? And so, you you need to build systems around around these things to make them not not you know, to make them functional, right? And the open-source world is experiencing that in particular, right? Where, you know, a whole bunch of projects suddenly appearing cuz somebody vibe coded
4:19 something they thought was cool overnight. A whole bunch of people are like, "Oh, I can fix my favorite project by modifying it to do something. I'm going to send in a patch." But they haven't actually reviewed it, or they don't even really they've never coded before. They don't even understand what this patch does, and they're the maintainers of the project are just like, "Look, there's there's a hundred of these. I don't have time." So, I mean, some people talk about like the the the hundreds of these things,
4:41 thousands of these things. GitHub is just exploding with what, you know, some people say is like AI slop, you know? What There's just so much out there, and then so therefore, you know, engineers and tech you know, technical people are just stuck going through it all, through it all. And then then you're stuck in this moment of sort of like what what was this all good for? I mean, how do you address that kind of that that that sort of issue of this sort of like having just this just a massive surplus
5:04 of of of product out there? Yeah. I think I mean, one of the things I observed that I think people don't necessarily realize is that AI is at least as good at reviewing code as it is at producing code, right? And there's kind of in the open-source world, with this AI slop, there's this interesting like it's a you can think of it as a giant distributed system of people producing the slop. Anybody can do it
5:27 and upload it to your GitHub repository and make your life as a maintainer miserable, right? But you can also get an AI to review those incoming PRs and tell you whether they're slop or not, and sort of automatically disqualify them if they're slop. But that would require the maintainers to set up some stuff, and And like expensive time that hasn't really happened very consistently yet. People are only just getting
5:49 started with this kind of automated review process, right? But you can do very good code reviews. In in the the like extended version of this, you can have automatic code reviews of PRs coming into your project, and then the person who's bought uploaded the the code can actually read the code review and then fix the code, and it can do this in a few cycles until eventually it gives up because it's always going to be slop or it's actually come up with a pretty good change, right? And once
6:13 you've got a pretty good change, then you should bring the human in to say like, "Hey, this is a pretty good change. Do you want it or not based on the vision for your project, based on your design principles, based on these other things?" And that's the opposite of the kind of exhaustion that we've been creating today, right? So it's all about the processes that you put stuff in. Um Bailey, we had this moment of the the Mythos moment very recently.
6:35 Um I just wanted to for you to sort of like in in in the work that you do, this has there really been this uh we've heard about the Mythos and and about, you know, Firefox like where you said they discovered like these these these um these these fallibilities. I was just wondering if you could just sort of are you in agreement that this is definitely a moment where there are new sort of security dangers out there um
7:01 that there weren't before? Yeah, I mean Mythos scared us all a little bit when they found vulnerabilities across Firefox, FreeBSD, and you know, a ton of other things. And you know, as AI models improve, they're going to find better vulnerabilities. There are some extremely complex and intric- uh intricate like exploits that can be done. And some which are so hard that, you know, humans
7:24 may not be able to do that. And while Mythos like they created a monster, really. Um and they decided like let's gatekeep this and let's try and like roll it out in the the most cautious uh way. But like this is AI. Anybody can innovate in it. Um Um, know, we have frontier models coming from American companies and then suddenly China comes out with DeepSeek and it's literally
7:46 open source. You can just download a state-of-the-art model. Now, what happens when China comes out with a Mythos alternative? Does that mean that now every single person that can, you know, self-host an open-source AI model now has the ability to hack into almost any system on the planet including some of the most secure and open-source um projects that exist?
8:10 That's a very good question which I'm going to put to [laughter] Well, uh I think I'm actually more optimistic than that. I think the timing is really important, right? There's the attackers, the red teams, there's the defenders, the blue teams, and there's what I would call the arms dealers who are selling tools to both sides, right? And I maybe they don't want to think of themselves as that way, but they they have to, you know, they built this incredibly
8:31 powerful tool that can be used for good or it can be used for evil. And I respect Anthropic a lot with Mythos in particular. Like they've been holding back on availability of this thing probably for multiple reasons probably including for marketing reasons, but it it it finds real serious security holes in real critical infrastructure software. And if they'd given that to the bad guys first, those bad guys would be exploiting that software before we
8:53 have a chance to fix it and we would be in really big trouble, right? The fact that they didn't do that, the fact that they're giving us a chance to like, hey, we can fix our stuff first is like is a gift. And then the question is like, well, okay, as the models keep getting better and better, are they just going to find more and more obscure security holes and it's never going to end in this like sort of downward spiral into doom? Or are we actually just going to like is stuff just going to get more
9:15 secure and not really hard to attack? And I think it's actually more like the latter, right? Like most software in the world today has never had a security review ever, right? And even if software that's been security reviewed gets reviewed once or twice a year, gets pen tested occasionally, some of the PRs get security reviewed cuz they're considered to be sensitive and some Some don't, right? So like most code has just never
9:37 been looked at. And now you have these tools that can look at it constantly. Every single change you make can look investigate and and try to find one of these security holes. And so, I believe that if we all get on board with caring about security, which is which is a stretch. But, if we do, we can secure our software. Uh getting people on board to to to worry about security may involve
10:00 regulation or may involve governments. Is that sort of something are you um Baylor, are you reluctant to see the government taking too much of a a role in this or is that something that you think there might be a way for them to to there there is a place for for regulation here? There might be a way. I think, you know, in answer to what you said, I don't have access to me those. Something exists out there that may or may not be able to completely break my
10:23 code base. I don't have a good opportunity to defend against it. And I think much like how, you know, the government has created structured and regulated markets which enable, you know, like each party in the the stock market to be able to to trade in a fair way. Um there might be opportunity for, as long as they don't overreach, um to try and like equalize this stuff.
10:46 That might be through export controls, that might be through um you know, laws and acts. But, uh yeah, I think there's definitely a little bit of uh equalizing that needs to happen. Um at the opening of the of the Web Summit um uh Paddy Cosgrave sort of talked about like these two worlds that, you know, there's the open source, you know, there's those who say that open source is already won and it
11:09 is the way. And then there's the others who saying that, you know, the the the US frontier models have have have already won and it's the way. Um the reality is somewhere in between there probably. But, I was just sort of wondering um on the sort of big sort of big picture level, every where do you sort of see things right now? Well, I think I mean, I think all the models have their place just as in any economic system. There's going to be the premium product and there's going to be
11:32 like all the levels down to the mass-produced product, right? And I think the most expensive tokens from the most expensive models are going to get more and more expensive. That's what I think will happen, cuz there's a shortage of GPUs, and like if you want this good stuff, you're going to have to bid for it with everybody else, and the price is going to get eye-wateringly high. You think it's high right now? I think it's going to get even more high. But at the same time, the price of like the tokens that are like a year behind
11:55 that are going to are going to crater, right? And that's that's going to be really exciting. Like pick the tokens we want. I think from a security point of view, that's going to be like if you're building infrastructure software that everybody depends on, the Firefoxes, the FreeBSDs, the Linux, the Chrome of the world, and I guess the Tailscales of the world, you're going to have to pay a lot of money for like the best security defense, right? Uh what I think about open source software, like, you know, a
12:19 lot of the stuff that we're all going to vibe code for ourselves or for the five people on our team, right? You're never going to be able to afford to run the best frontier security software, for example, to review that. You have to do something else, right? My my proposal is don't put it on the public internet. It doesn't need to be there, and if it's not on the public internet, nobody can attack it, right? That that's the best approach unless you're building these
12:42 fundamental front-end-facing public products. Um Bailey, token maximization, you know, we hear these stories about, you know, like at the big corporations or people are told to spend as much as possible, you know, what is it spend? To use as many tokens as possible, which ultimately means spending. I'm I was just sort of wondering, how does a company like yours approach like, you know, token you know, just using up of tokens? Is it like
13:05 go crazy, or like how do you try I mean, how do you try to how do you navigate between like showing that you're utilizing these tools to their fullest potential, but at the same time, you know, staying afloat, you know? I mean Yeah, I'm a little bit against the whole like token maxing hype train. I think, you know, a lot of things are hype trains in this industry. Even to some degree, people want to be like open source just because it's the popular
13:27 thing or not because it isn't. And it's really about like using what's sensible for you. Now, I do completely agree with Avery on the whole thing about being a little bit sensible and a little bit cautious about AI adoption. Um, do I think that like everybody should be spending as many tokens writing as many lines of code as possible? No. We've known for years that like lines of code does not equal output. Um,
13:51 and so I think really you want people to be AI-enabled. I can do a better job with AI. Everybody on my team can do a better job with AI. But, you know, otherwise like let's keep it sensible. Right. Uh, Avery, you know, to people in the audience, um, you know, people that are, you know, coding, who are coding for a living or thinking about coding for a living, um, what is sort of your advice to them in terms of, you know, being in
14:17 this world where, you know, some people say AI and they think it's a magic wand and then something magically just appears and it's fit for purpose. Uh, there's still a a major role for, you know, the good old human being in the story, isn't there? Yeah. Well, rather than a magic wand, maybe I I can compare it to a genie, where you make a wish and you literally get what you ask for. Uh, and it turns out not to be what you wanted, right? Cuz that that's what
14:39 happens over and over again. It's like it creates that it is magical. I get it's magical in the most literal sense of like we don't even know how it works. Even the people making it don't know how it works. And it and it grants wishes, but it grants them in like strange ways that you might regret, right? Uh, and I but I think, you know, it's a tool just like just like a human is a tool. Humans are magical, right? You ask them to do something and you don't always get exactly what you wanted. And sometimes
15:01 it's good and sometimes it's bad. And we've got tens of thousands of years of building human society around the fact that humans are magical, right? But I think, you know, society exists to serve humans, right? It doesn't exist to serve computers in a like running mathematically calculations that simulate humans, right? And so it it's up to us to do what we want with the tools that we we
15:25 have, right? And these these tools can be used for anything, right? They can be used for attacking, they can be used for defending, but they can be used for like writing a bunch of slop code, but they can be used for defending and fixing slop code without humans having to be involved. So then we have to up-level ourselves to this like I'm going to think about the abstract. Like okay, I actually have this you know, this this patch came in and maybe a bot wrote it, maybe a human
15:48 wrote it, who knows, right? But like it works now. It's I've verified that it does what it's supposed to. I've even verified it doesn't have security holes. Is it what I want to exist in the world? That's what an open source maintainer will have to decide, right? And that's the that's the fun part of being an open source maintainer. That's why we get into it, right? When once we start, we realize it's like even before AI, it's like 90% like jerks writing
16:12 angry posts in your issue tracker or like triaging stuff or like slop that might or might not have been AI generated PRs, right? But like the 10% that was fun is like working with people, building something cool, and having someone send in something that makes your cool thing even more cool, right? And that can be the whole job because we can eliminate the tedious part using these tools. But it's the human's job to decide what they want to exist in the world. That's what we all
16:35 work and be are going to be able to do with these more advanced tools. Billy, do you have I mean what what is your perspective on that too? I mean in terms of like, you know, you may not be like the most senior, most most experienced like, you know, programmer, but suddenly these vibe coding tools like allow that person to to to explore ways that they they couldn't have done before. Yeah, I know for us we said we're never going to fire anybody
16:57 because of AI. I mean we never over hired the team in the first place, but now we just expect people to be able to to produce more. Um you know, like I said, this doesn't have to be going crazy at it and you know, uh overusing tokens just for the hell of it, but you know, it's it's a very exciting time to be in open source, to be in anything, really. The world is rapidly changing each week
17:21 to the next. We have more and more capability. And, you know, overall, this should be able to be a positive that we can all use just to build more things and, you know, invent. All right, cool. I think we have to we have to leave it there. Thank you very much, guys. That was really cool. Thanks a lot. Thank you. Thank you.

