This looks like measuring what is easy to do, rather than what really matters.
PR open counts, issues created , ceos/founders spending more time on linear don't automatically lead to better outcomes (in my experience they are often negatively correlated:-) )
This looks like a "We are so AI native and efficient!!" article. At least, they could delve deeper into how they define the metrics and how they collected the data.
this seems inappropriate. I think its a bad paradigm that just because you use a platform's service, they get intimate details about your usage. And for them to be so bold about publishing the statistics they've stolen from their customers data? Gives me a reason to never recommend my org use this platform.
The data is aggregated, and you cannot possibly identify a single user from what's been published. I see no issue here.
I think their point is that there's business value in the usage data and linear is using that value in a way that benefits them but not the customers they got it from.
It reminds me of matt levine's reframing of insider trading where it's not about fairness it's about theft. You're supposed to get secret insights and use them to get an edge. What you can't do is get an edge for yourself with secret insights that your employer got.
So it roughly comes down to "that data is valuable and rightfully belongs to the originating company." Which then makes this a contract diligence type situation.
My company uses Linear. The data presented in this blog post is worthless to me and the company. It can be crudely summarized as “agents and agentic development processes are conducting more Linear operations.” Duh!
lol isn't this what google said and yet..the NSA cometh
I’d rather it was published for free in a blog post than sold without my knowledge. As a linear user. This shit doesn’t matter, everyone in the industry know these kinds of stats are being tracked, I’m happy for a company to be transparent about it.
Why would a company want to leak its processes and workflows to another company in a capitalist system. Seems sloppy and a short sighted transfer of wealth to external stockowners.
Uhh that's how a lot of economic data works too. Guess how we get a lot of jobs data... ADP
Oh hey! My industry. Guess what? ADP doesnt just yoink your data. ADP conducts voluntary surveys on the scale of hundreds of thousands. Voluntary. ADP also pays for it for the most part. Do you think Linear conducted..voluntary surveys here?
This is pretty interesting but I wish it would have been refined in two ways:
- The prose before the data appears to be AI generated. Not a big deal but it makes the reader work harder to figure out what's actually being said.
- Linear didn't control for their platforms AI changes over the past year. The platform has become much more AI integrated, so a lot of numbers will move. I'm not sure how you do it but this is only useful signal with a control.
I think it's worth noting that headlines like "AI adoption has spread to every function" is only limited to roles covered/tracked by Linear.
Recent studies by e.g. Google show a much broader range of adoption.
These days, my work has become: generate code for 20 minutes, then spend an hour reading it.
And then more hours cleaning it up and re-prompting.
I program as a hobby, personal projects because I can.
I recently set up a local llm to see what the fuss is about and other than the few ringer solutions my experience is as you described. 2min promping, 5min waiting, 3hrs debugging or just doing it myself.
I am very likely doing it wrong, and it does speed up some aspects, but I wouldn't say I trust llm code any more than my own. Until it runs and throws an error, the llm is 100% confident that it has written perfect code.
Models that run on (average) consumer hardware are not even close to comparable to models like Fable or Sol. Its like comparing an ant to the largest dinosaur.
A normal agentic loop will have the agent using a type system and basic tests to do some basic validation of changes. A good agentic loop would give the agent a very easy way to verify if it’s on the right track. I think agents are better than many humans at writing error free code (runtime errors, not bugs. The code could still be buggy or incorrect.)
If you built a task management system, you'd have very different code bases depending on whether it's for internal use at a mid-size development org or as a SaaS.
So I wonder whether, in your experience, the results you've seen, could have improved by providing sufficient context? - or what context was given.
I.e. if you have the agent that same context, as one of your colleagues would have/require to solve a problem.
Try using Fable and report back. Local LLM is to Fable as Little Tike car is to a Porsche.
I don't think setting up a local LLM is a reasonable way to get a good idea of how enterprises are using this stuff.
Local llms aren’t super exciting unfortunately.
Why not just write the code yourself? To me it seems like methodically constructing the pull request by hand is probably faster than lazily prompting and re-prompting.
What I've observed is that by prompting for longer I get to keep my brain focused on the architectural ideas (networking, protocols, data structures, etc) rather than worrying about the most performant/elegant implementations.
I always enjoyed writing code and I am very good at writing very performant and elegant code, but it would sometimes come at the detriment of focusing on the code and not the design.
because my manager will ping me and say "anon you aren't prompting enough" like they never heard of Goodhart's law before.
No, don’t be silly, they get ai to spy on you en masse instead. Nobody has to look at anything anymore for it to be actionable.
In my experience?
It used to be that one person had one to three codebases they knew intensely at my company. If you needed a bug in codebase X fixed, person Y was the one to do it and if they aren't available, person Z can do it, just not as quickly.
Now every person on my team has to handle tickets for every single codebase. There are about two dozen different large codebases involved here.
It's a ludicrous antipattern because person Y still needs to review the PR that person A generated for codebase X, and it will take them about as much time to wrangle the 2000 line PR (oh boy do LLMs love their mocks for unit tests) as it would have been for them to do the 50 line code change.
On top of that, it has "allowed" us to add feature after feature onto codebases not designed for them without refactoring. Is it good that this Flask API went from a purpose built service that interacted with the data analytics for product A stored in database X, and now our sales guys can sell product B, C, and D stored in database X? Uh, I'm sure it's great for them. Oh, and now it all can be stored in database X, Y, or Z depending on what the customer wants or what sales promised them. Great. Now I'm looking at a Flask app.py that's 12,000 lines of repeated code.
It has allowed poor designs to still produce working code. For a while. We seem to be getting a lot of bugs lately that look really bad to customers because it's for really simple shit. And I can't help but notice that happening to all the various products and sites I use too...
Companies are tracking token usage across the board. You are in trouble for too less or too many. Unfortunately the token usage is the only measurable thing for most folks so everyone is playing the game, otherwise how else would Anthropic and OpenAI make the money
Writing code directly takes longer to warm up. Usually, I'd keep tens of thousands of lines in my head. In the past, I spent a lot of time designing error propagation and execution contexts. (Talented people might figure it out right away, but unfortunately I don't have that kind of talent.) So I'd have to think about things like Result<T> and how far to propagate errors—and worry about whether my approach would conflict with the existing codebase.
But these days, AI just generates code following the existing patterns of the codebase. In the past, staring at a blank screen meant going through a checklist of things to design—starting from policies and writing everything down step by step. Now, I just ask AI and it gives me a template—which is great. Then if the AI makes a mistake, I fix it manually.
Of course, I still hand-code sometimes—but only in the areas I enjoy. Most of the time, I use AI coding. Both are fun, and they complement each other in interesting ways. Doing both together is actually enjoyable.
It does so better the more "standard" the "existing patterns" are. :^)
I feel similarly, but at the same time, I think I am the exact opposite. I actually find formulating hypotheses more fun.
For hobby projects or things I start casually, I usually do not think about errors and such at all. When it is a tool I want to build or need for myself, I really do not care about that part.
In my case, I do not contribute to open source at all. Mostly, I deliver code for factory systems or specific companies, and usually, there are strict enterprise requirements. (To be precise, there is always that mandatory code the lead developer on their end dictates, right?) That kind of code is mostly no fun, but it has to meet their requirements and often clashes with my own style. Having AI write that code for me is a huge relief.
In that sense, I think it is just a difference in personality and preferences. I originally became a programmer because I wanted to make games. I started programming because I found it fascinating to see things drawn and displayed on the screen. Becoming a programmer was all because making Flash games was so much fun... So in that regard, for me, writing code is just 'drawing what I want on the screen', which is why I guess I do not mind if the code is written by AI.
When I contribute to other people's projects, I do not use AI for anything other than English translation, but for my own projects, I have no hesitation.
Is this really just a difference in inclination? It is not that I did not enjoy writing code, but rather that seeing what I want rendered on the screen brings me more joy.
When the concept of 'vibe coding' first came out, I really hated it (since my knowledge was earned over 4 to 5 years of getting scolded by lead developers as a subcontractor and factory software provider). But thinking about it, what I really wanted to do as a developer was just to build the worlds I envisioned, so I decided not to let it bother me too much.
We talk often here on HN, and I really enjoy debating with you. I learn a lot from you.Mr."skydhash", I actually remember you quite often, and I even steal a few keywords from your posts sometimes. Because we have different tendencies, we occasionally clash, but having these conversations is exactly what makes it enjoyable.
Thank you for always replying. Have a great day, and I hope this does not offend you in any way.
Felt that. I already gave DeepSeek the HTML and it still said the UI was good to go, then told me it has no vision.
Like the Titan submarine team's moto mine is - real men test in production.
The LLM produce so much code that the best I can is skim and look for obvious flaws, also pass it trough adversarial one.
I spend 45 minutes writing it. then i git commit and move on cuz i made something good.
who is winning here? lol
[dead]
That's gotta be at least 10x or 20x more efficient than the old way of doing things.
It could be. Or it could be 1x, or 0.2x. You don’t have enough information to make that judgment.
I was joking. The workflow doesn't seem particularly fast or engaging in my opinion. AI code generation doesn't seem worthwhile to me.
> I think it’s near time we all stop having such strong opinions about the matter either way personally.
I don't think this is reasonable given how abusive the pro-AI rhetoric has been for years now.
I’ve reached a similar conclusion, but there’s a part that worries me: the expertise that allows us to judge AI’s output was itself built by doing the work we’re now delegating. So there’s a risk that our judgement will decay over time. I’ve been thinking about the problem as choosing where we can afford to “borrow” comprehension, versus where we need to keep exercising it, and how to “claim back” the critical comprehension we lost.
That plus if you’ve ever stared at your code and then searched StackOverflow to see if you could find a better way of doing it, it’s like having that running continuously.
Nice data!
Tim (author) if you're there: it'd be amazing to see the split of which agents people are using, if you have that data.
I didn't know Linear has "AI features". Linear is boring, but that's actually fine by me.
I use LLMs to write my code, but this does not show up in this data.
"Pull requests are up 111% in two years". Would be more honest to say that the number of pull requests "detected" by linear are up XXX%. Because it only works if you setup git repo tracking and use it properly. And at that point it is not obvious if more teams are using linear and using it correctly, or if the number of PR really increased that much!
The AI Slopologists strike again. More garbage by garbage people.
Is this just an opinion? If so, fair.
If it's an attempt at rebutting their claims etc, it'd be easier to interact if you provided some data, or concrete observations :)
Can't tell if you mean the people the article is talking about or the article itself
¿Por qué no los dos?
> Time spent on customer requests, docs, and projects held steady [..] AI has so far changed how teams execute far more than how they decide what to build
I think the measurement for this may be flawed. We do mostly use AI to decide how to build. But what we build is influenced by AI-driven research into a problem or task. That's largely done in coding and desktop AI tools, not Linear Asks/AI.
I'm working on accelerating my team's work by implementing AI-driven code pipelines with guardrails to eliminate as much unnecessary review time as possible. Also making a chatbot for turning repetitive tasks & PRs into buttons, and an "architectural guidance" chatbot that gives advice tailored to our business, software/system architecture, cloud, standards, etc. This puts AI and automated jobs in the center of both how (automated task) and what (architecture guidance).
But this has a not-so-great implication for Linear. With my tools, a human never has to touch a ticket, so we could use any ticketing system with an API or CLI. Linear is a great product because they made a great interface. What happens when I replace their interface with a chat bot?
[dead]
[dead]
[flagged]
[flagged]
[dead]
[flagged]
[dead]
This looks like measuring what is easy to do, rather than what really matters.
Even that's hard. There aren't enough signals to attribute changes directly to AI, so these apps seem to correlate the signal that the user was interacting with AI to the changes they made e.g "Bob used AI at that time, and they opened a PR at a similar time, so Bob probably used AI to make that PR."
Until the tooling for gathering data on AI usage improves the data will be fairly interesting because correlations often point to something related, but won't be a source of truth.
Yeah all of these along with raw token usage are metrics that were being used around December to February by people who just didn't have anything to go off yet.
Skill/Hook usage rates, budget spend, auto-approval rate, focus area heatmaps, MTTR, MTTD are all there now
I'd be interested to overlay, I don't know, customer satisfaction or anything that can show the follow-on effect of all this output. Linear won't have that information.
My guess is some will jump up (where the team has managed to make themselves move effective and responsive) and many will plummet (doesn't need explaining).
Then there might be something to look at.
I'd be interested to overlay, I don't know, customer satisfaction or anything that can show the follow-on effect of all this output. Linear won't have that information.
Very few businesses can accurately attribute customer value to the work they do, especially once they're passed start-up scale. A mature company makes lots of small changes and they're rarely measurable.