Energy, data, environment
1494 stories
·
2 followers

Where are the AI productivity gains? - Bennett School of Public Policy

1 Share

Outside software development, the evidence that AI is improving productivity remains thin. The barriers are less technical than organisational, argues Sam Gilbert.

Huge investments are being made on the assumption that generative AI is about to deliver transformative gains in economic productivity. Between them, Microsoft, Meta, Amazon and Google are expected to spend more than $600bn on AI infrastructure in 2026, while global capital expenditure on advanced chips, data centres and power generation looks set to exceed $1 trillion – more than the capex of the entire oil and gas industry. Analysis by The Economist suggests AI is on track to be the largest investment boom in history – bigger than the dotcom boom; bigger even than the railway and canal booms in the 19th century. But to what extent are these investments justified by real evidence of productivity improvements?

AI adoption is broad but shallow

ChatGPT now has more than one billion active users, a milestone reached faster than Instagram or TikTok. But striking adoption statistics like these can obscure how much most people are using AI tools. Gallup polling suggests that 52% of US employees use AI at work, but only 15% use it every day. Similarly, a National Bureau of Economic Research survey of ~6,000 executives in the US, UK, Germany and Australia found that while more than two thirds use AI, they only use it for an average of 1.5 hours per week.

Analysis of companies’ spending on AI tells a similar story: the payments platform Ramp, which processes transactions for around 70,000 businesses, reported in July 2026 that while the top 1% of firms spend roughly $7,500 per employee per month on AI, the median spend is only $12 per employee per month.

Coding is the “killer app”

Software development is the one domain where the productivity case for AI is more or less proven. More than half of all new code is now AI-generated, and analysis by the venture capital firma16z indicates AI companies’ coding-related revenues are more than double everything else combined. Academic working studies of AI coding tools like Cursor and Codex consistently show productivity improvements in the range of 20-40% (Cui et al 2025, Agarwal et al 2026, Sarkar 2026, Becker 2026).

But even here, the gains are contingent. Drawing on commit data from around 100,000 developers, research by Yegor Denisov-Blanch and colleagues at Stanford finds that the biggest productivity benefits come from low-complexity tasks and new projects using popular programming languages such as Python. For complex tasks on legacy codebases in less common languages like COBOL, the gains from using AI shrink dramatically or even turn negative. As public sector technology development is predominantly “brownfield”, this is an important finding for policymakers to take on board.

The better the evaluation, the smaller the effect size

Outside coding, most deployments of generative AI that have been independently evaluated involve AI assistants – either public-facing chatbots or internal-facing “copilots”. In general, the more rigorously these are studied, the less impressive the results look.

The UK Civil Service is a case in point. A cross-government trial of Microsoft Copilot, involving 20,000 civil servants, claimed average time-savings of 26 minutes per day based on self-reporting. But a subsequent controlled study in the Department of Work and Pensions found time-savings were more modest, while the Department for Business and Trade’s evaluation found no evidence of improved productivity, with some tasks slowed by the need to correct poor-quality outputs. This resonates with the Bennett School finding that using AI chatbots does not make workers more productive in and of itself.

In North America, the Canada Revenue Agency reported that its chatbot “Charlie” gave accurate answers to user queries in 70-90% of cases in internal testing, but an external evaluation by the Auditor General found it gave correct answers only a third of the time. Meanwhile, New York City withdrew its own citizen-facing chatbot after investigative journalists demonstrated that it was systematically giving unlawful advice to landlords about tenants’ rights. Little wonder, then, that according to KPMG only 7% of business leaders are currently reporting a positive return on their AI investments.

The most encouraging case is instructive about why. Caddy, an internal copilot built by the UK Citizens Advice with the Incubator for AI, was evaluated through a randomised controlled trial and produced positive results: response times to citizen queries were halved, with advisers much more likely to report high confidence in the advice they gave.

Rather than seeking to replace advisers, Caddy was specifically designed to help them become more productive. It may be that in customer- and user-support contexts, realising productivity gains is dependent on a “human in the loop”. This would be consistent with research showing that AI is not a “plug and play” technology, but depends on strong management practices to deliver results.

Rising costs and externalities

Another factor that may limit productivity gains from AI is that the scope for cost-saving opportunities is sometimes smaller than assumed. In functions like insurance claims, much cost has already been stripped out through decades of business process optimisation, outsourcing, and offshoring, meaning there is not much left for AI to unlock. At the same time, AI services are getting more expensive: having subsidised usage costs to drive growth between 2022 and 2025, investors in OpenAI and Anthropic now expect margin improvements from higher prices as the companies prepare to go public in 2026.

Even where productivity gains from AI are real, they are partly offset by externalities associated with the broad adoption of AI tools. Employees who propagate “workslop” – AI-generated reports, presentations and emails, that appear polished but lack substance – create an extra burden for their colleagues. According to the Harvard Business Review, more than four in ten workers report having encountered workslop, estimating that each instance costs them nearly two hours in rework. Meanwhile, a recent survey found that almost two-thirds of professionals have uploaded work-related information to public AI tools, exposing their employers to material data protection and commercial confidentiality-related risks.

There is also a demand-side externality of AI adoption with serious consequences for the public sector. Chris Schmitz and colleagues at the Hertie School have documented 84 cases of what they call “service flooding”, where AI-assisted applications, appeals and objections overwhelm the operational capacity of support functions.

Policy implications

To paraphrase American economist Robert Solow, AI is everywhere except in the productivity statistics. There are surprisingly few instances of AI-related productivity gains being realized outside software development, while evidence of negative externalities is beginning to emerge. This should moderate policymakers’ expectations about the potential for AI to relieve pressure on public services and improve national productivity – and about the timeframe in which this can be expected to happen. Historically, major technological breakthroughs have taken decades to transform the economy, and as yet, there are few reasons to believe AI is a special case.


The views and opinions expressed in this post are those of the author(s) and not necessarily those of the Bennett School of Public Policy.

Authors

Related Articles

Blog

Boy wearing white shirt and black shorts carrying backpack standing on black concrete road between vehicles and trees during daytime photo – Unsplash

Blog

woman fixing cables in data centre

Blog

computer screen with cursor arrows and hands

Blog

Blog

News

Podcast

Blog
Read the whole story
strugk
7 hours ago
reply
Cambridge, London, Warsaw, Gdynia
Share this story
Delete

Floating data center would provide water and electricity to 32,000 homes

1 Share

Some companies build a desalination plant and call it a day; others build a data center, while others, still, build power generation facilities. One project wants to build all three … into one facility … offshore … on a movable ship.

Introducing the Blue Economy VITAL 100 MW Kraaken, a proposed infrastructure platform developed by marine engineering firm InMar Technologies and energy systems developer OptiFuel Systems under their Optimal Transit maritime technology partnership.

The concept combines offshore power generation, a desalination facility, and an AI data center into one large floating vessel that can sail away when required. According to Optimal Transit, an onboard plant would generate up to 40 MW of continuous electricity, which would be sent ashore, enough to power roughly 32,000 homes 24/7. Another system would produce 30 million liters (7.9 million gal) of fresh water every day, enough for an estimated 150,000 people. On top of all these, the vessel would retain 60 MW of capacity for AI-grade data-center computing.

It would require no land, no grid connection, and, according to Optimal Transit, no conventional fuel for those operations.

The Blue Economy VITAL 100 MW Kraaken is the latest configuration of a concept Optimal Transit unveiled in July, primarily as a self-powered floating AI data center, with 10-, 20-, 50-, and 100-MW variants. Virtually all of its available output was intended to be absorbed by racks of GPUs. The new VITAL configuration uses the same basic 100-MW power system, hull, and mooring infrastructure, but splits the output between compute and utility services for nearby coastal habitats or disaster areas.

Now, you probably have a lot of questions about this wonder project, but let's start with the most obvious. Where is the proposed 100 MW of electricity supposed to come from?

At the heart of both versions is Optimal Transit's patented Digital Ocean Thermal (DOT) engine. Fundamentally, DOT is an extensively modified Ocean Thermal Energy Conversion (OTEC) system. OTEC exploits the temperature difference between warm surface seawater and much colder water hundreds of meters below. In a closed-cycle system, the warm water heats a low-boiling-point working fluid, such as ammonia, until it vaporizes. The vapor drives a turbine that generates electricity. Deep cold seawater condenses it back into a liquid, and the cycle starts again.

Kraaken adds an interesting second heat source: its own computers, repurposing the waste heat that other offshore data center concepts, such as China's “world's first” underwater data center, release into the surrounding ocean.

A diagram of Optimal Transit's OTEC process

A diagram of Optimal Transit's OTEC process

Optimal Transit

Optimal Transit's published diagram shows warm seawater entering at around 25 °C (77 °F) and heating ammonia in an evaporator. Meanwhile, liquid-cooled servers aboard the ship send their waste heat, at roughly 45 °C (113 °F), through an integrated thermal bus to what the company calls an AHEB, which further heats, or "supercharges," the ammonia vapor before it reaches the turbine. After expansion, 5 °C (41 °F) deep seawater cools the ammonia in a condenser, and a pump returns the liquid ammonia.

The company additionally refers to a multi-stage Rankine cycle, enthalpy recovery, supercharging, and green-ammonia synthesis as parts of DOT. More specifically, its diagram claims the architecture can reduce turbine size by around 70%, cut the size of the cold-water pumps and condenser by 70%, reduce warm-water pumping and evaporator costs by 40%, ammonia pumping costs by 40%, and cold-water pipe size by 70%. Impressive, albeit yet-to-be-demonstrated thermoelectric power generation.

What is particularly clever about the Blue Economy VITAL is that the same thermal infrastructure also performs other heavyweight tasks, including desalination.

It would produce fresh water using vacuum-flash desalination. At sufficiently low pressure, warm seawater boils at a temperature far below the usual 100 °C (212 °F). The resulting vapor leaves its salt behind and can then be condensed into fresh water using the available cold-water stream. It is an established concept closely related to open-cycle OTEC, in which fresh water can be produced as part of the thermal process.

Kraaken therefore attempts to squeeze electricity, cooling, and fresh water from essentially the same hot-and-cold thermal infrastructure.

Now, you would expect such a facility to be anchored in place to the seafloor. Not the Kraaken. It’s itself a ship that uses a Small Waterplane Area Twin Hull (SWATH) design. Rather than putting most of its buoyancy at the wave-tossed surface, a SWATH vessel places much of it in two submerged hulls connected to the upper structure by relatively narrow struts.

The original 100-MW configuration of Kraaken

The original 100-MW configuration of Kraaken

Optimal Transit

The original 100-MW Kraaken concept is a roughly 300-ft (91-m), 50,000-long-ton vessel, while the smaller 10/20-MW design measures around 250 ft (76 m). The modular data-center hardware is liquid-cooled and replaceable, allowing newer generations of GPUs and servers to move aboard without replacing the long-life ship underneath. This is one reason the creators call the ship scalable.

So, how do the promised utilities, electricity and fresh water, as well as the data connection, reach land? The entire system is connected to the shore through a quick-disconnect umbilical system, “quick-disconnect” being a particularly interesting word. Should a major storm threaten, Optimal Transit says Kraaken could detach from its mooring within hours, move away under its own propulsion, and return to reconnect once conditions improve. The earlier design specifies speeds approaching 16 knots (18 mph/30 km/h). It could similarly leave altogether if its contract ends, political conditions change, or a disaster zone needs it more.

Optimal Transit envisions deploying the vessels off islands and in poorly served coastal regions, industrial areas, and disaster zones. It also proposes clustering five VITAL vessels within a 2-mile (3-km) offshore area to create what it calls a "Sovereign Power Park." Such a cluster is claimed to provide around 200 MW of electricity to shore, 40 million gallons (151 million liters) of fresh water per day, and 300 MW of data-center capacity.

Next question: how much would the whole thing cost? Optimal Transit puts an all-in VITAL vessel at about US$587 million, roughly the same amount as or less than the price of Jeff Bezos’s data-center-less, non-freshwater-producing luxury yacht. The developers estimate that separately building a 40-MW power station, a 7.9-million-gallon-per-day desalination plant, and a 60-MW AI data-center shell on land would cost a combined $750 million to $1.33 billion. By the company's calculations, that makes Kraaken about 44% to 78% of the cost, while taking around three years to deploy rather than six to 10.

Now for the reality checks. First, there is an important distinction between Kraaken's components being based on established technology and Kraaken itself being established technology. There's currently no 100-MW Kraaken bobbing offshore, powering 32,000 homes, at least for now. In its July announcement, Optimal Transit said Series A funding would fund ABS-ready engineering drawings and comprehensive digital-twin validation of DOT. Production plans depend on a proposed Series B raise in 2027.

Secondly, 100 MW is a formidable number for ocean thermal power. While the physics of OTEC is well understood, the small ocean temperature difference makes it inherently low-efficiency. A 2026 study modeling a 100-MW-net OTEC plant obtained a thermal efficiency of just 3.75% at its best-performing 700-m cold-water depth. This translates into astonishing quantities of seawater. The US National Oceanic and Atmospheric Administration (NOAA) estimates that a conventional 100-MW OTEC plant could move 10 to 20 billion gallons (38–76 billion liters) of seawater every day, with a cold-water pipe around 33 ft (10 m) across reaching approximately 3,300 ft (1,000 m) deep. The deployment of that pipe remains one of OTEC's major engineering challenges.

Optimal Transit specifically claims DOT can slash that cold-water pipe by 70%, which, if demonstrated, would be a very big deal. However, it also raises an obvious question for a vessel designed to disconnect and sail away from a hurricane. Exactly how does its deep-water intake connect, disconnect, and survive that process? The public material doesn't yet provide enough engineering detail to answer it.

There are smaller accounting questions, too. The advertised 40 MW of exported electricity plus 60 MW for compute already add up to the full 100 MW, while desalination, seawater pumping, and the ship's own auxiliary systems also need energy. The company describes the figure in terms of net output but hasn't publicly provided a detailed power balance showing exactly how it accounts for all those parasitic loads.

Optimal Transit also says DOT could operate from equatorial to Arctic waters. Traditional OTEC is overwhelmingly associated with tropical waters because NOAA says a year-round temperature difference of more than 20 °C (36 °F) is desirable. How DOT maintains useful output where that ocean temperature gradient doesn't exist is another detail we'd like to see explained.

Now, none of these limitations make Kraaken impossible. Its technologies – Rankine-cycle power generation, ammonia working fluids, waste-heat recovery, vacuum desalination, SWATH vessels, liquid-cooled data centers and offshore umbilicals – are hardly science fiction. However, the extraordinary part is putting it all together, delivering 100 MW net, fitting the system aboard a movable ship, and doing it for $587 million.

Source: Optimal Transit

Read the whole story
strugk
17 days ago
reply
Cambridge, London, Warsaw, Gdynia
Share this story
Delete

Critically Evaluating LLMs: What Data Visualization Can Teach Us - Nightingale

1 Share

These thoughts are gathered from the before, during, and after of a workshop I led as part of a data visualization conference. Big thanks to the Bar Chart Club Conference, led by Erin Waldron, for hosting the original workshop and helping to make it a success.

Introduction

In March I led a three-hour workshop called “Adding AI to Your Data Visualization Workflow.” A more accurate title might have been “Practical Tips for Thoughtful Engagement with AI and How to Resist the Creation of Meaningless Slop.”

We’re now on year four of large-scale public access to large language models (LLMs). The rosiest corporate-backed shine is definitely starting to wear off of generative AI (RIP tokenmaxxing, for example) as questions about financial cost and ROI become more urgent. While commonsense financial reservations are finally gaining more traction, generative AI is likely here to stay, at least in some capacity. 

Whether LLMs stick around or not, evaluating them through the lens of data visualization has changed the way I interact daily with these and other tech products.

Over the last four years, I’ve worked as a curriculum developer for Codecademy (now owned by Skillsoft), designing online interactive content in data visualization and analysis. Like many people in tech or tech-adjacent roles, I’ve been told multiple times to use AI whenever I can. Looking at AI through this data visualization lens is helpful and fairly comprehensive, because it requires attention to text, numbers, images, and code. Data visualization was my AI testing ground, and forced adoption by my employer kept me at it when on a personal level I would have said no again and again. 

The two major challenges that I ran into repeatedly are these:

  • Generative language models will always have the potential to hallucinate. They regularly return untrue statements, and this problem is inherent to how they function.
  • The LLMs most of us interact with are controlled by companies. The built-in sycophancy of chatbots to which you are, first and foremost, a customer is antithetical to transparency. We require transparency for data work.

If we choose to (or must) use the technology for data visualization work, how do we mitigate these harms?

There are a lot of things we can do. My solutions focus mainly on the ways that we think and talk about LLMs and how those translate into critical use. The solutions I propose here are questions for critical evaluation at multiple stages of projects (e.g., before prompting or afterward while evaluating a result). They fall into three major categories for thoughtful, effective use of AI:

  • Accuracy: Is it true and verified?
  • Security: Is my data safe and private?
  • Sovereignty: Am I using the tool the way I want to?

While digital sovereignty can refer to laws and (inter)national-level regulations, in this case I’m talking about the choices we make about the tech we engage with. Where can we exercise control and self-determination and use our own judgments and values to guide our actions? While that’s often not part of the conversation because it’s the hardest one to quantify and probably takes the most introspection and critical thought, I think it’s maybe the most important as we consider what it will look like to exist alongside AI tools in the future. They’re part of the future, but they’re not the only part of it. 

So what are the questions?

Accuracy:

  • Am I asking the AI to scrape web data and give it back to me in some capacity?
  • Am I working within a defined universe (I’ll include the data and an example) or a more open universe? (I’m asking for something, but I’m not really sure what I want.)
  • Have I set up boundaries against sycophancy to improve my ability to interpret the result?
  • Can I verify the results? If so, did I verify the results?
  • Do I know enough about this topic to be confident that the answers make sense? 
  • One step further: Do I know enough about this topic to be confident that the answers are smart?

Security:

  • Am I working with private data? (Mine or someone else’s.)
  • Am I working with unpublished intellectual property? (Mine or someone else’s.)

Sovereignty:

  • Am I working on a part of the process I actually love to do?
  • In getting these answers via AI, will I lose an opportunity I value to find the answers on my own?
  • Is my question or idea worth the resources I’ll be using to explore it this way? (This question deserves its own article. Environmental degradation in service of AI tools is reason enough on its own to be a conscientious objector.)
  • Will I be comfortable disclosing that I used AI for this part?
  • Does it add value to my project? (Especially with text.)
  • Is it up to my standards?
  • Am I erasing my own voice or opinions by using AI for this?

I arrived at these questions through grappling with those two challenges I mentioned earlier: hallucination and lack of transparency. Read on for lots more context.

Part 1 of the problem: Generative language models will always hallucinate

I’ll start with what is likely a familiar exchange, from a current Claude model in March 2026. I asked for compiled data from a Wikipedia table. This is a task I do fairly often for simple, public datasets that are regularly scrutinized for accuracy, so I wanted to see if Claude could do it instead. It worked for a bit and then returned an impressively well-formatted artifact, a spreadsheet with color coding and helpful filters applied.

Screenshot of a chat with Claude. The user asks it to compile the data in the tables of a Wikipedia page, "List of power stations in Maine," into a single table with a column for type of power. Claude reports fetching the page, then says the web fetch gave it a partial view and it will try a more detailed extraction. After searching the web and viewing a file, it says it now has what it needs, that the web fetch gave it enough info, and that it will compile the data it knows from the Wikipedia page into a single table as an Excel file. It then presents a compiled spreadsheet of roughly eighty power stations with a color-coded Power Type column, beginning with natural gas, six stations.Claude compiling a Wikipedia table of Maine power stations into a single spreadsheet, March 2026. Screenshot by Eva Sibinga.

I noticed that longitude and latitude had been stripped from the output, so I asked why they had been dropped. Claude replied that the web fetch had likely not rendered the data correctly, so it dropped the affected columns and didn’t flag the omission. Okay, not a huge deal. 

I dug a little deeper into the table and realized that the data was completely wrong. Names of power stations had been changed slightly, but most importantly, the key metric I wanted (power in megawatts) was incorrect more often than it was right. I asked Claude to give me some insight into what had happened.

Screenshot of the same conversation continued. The user says much of the data in the table is wrong — plant names differ from the original and capacity in megawatts is incorrect more often than it is right — and asks for insight into why, without redoing the work or apologizing. Claude answers that the Wikipedia page's tables did not render properly when fetched, that the raw text output contained only column headers and a handful of fragments with no row data, and that rather than saying the coordinates and most of the table contents were not coming through, it filled in the data from its training knowledge, which is where things went wrong.Claude explaining why the compiled data was wrong, March 2026. Screenshot by Eva Sibinga.

It had failed to fetch my data and made up the results from its training data, then presented the table as a finished artifact rather than flagging that there was any issue. It also literally said “the web fetch gave me enough info” before making up the data.

Claude is also capable of doing this task correctly, which makes the failures much harder to catch. I’m not sure why the web fetch failed in this case, but I asked Claude to repeat the same task in the exact same language (in a different chat in the same account and on a different account) and it was able to give me a table with correct data. 

The issue of inconsistency means the tool is impossible to use efficiently. It appears to fail randomly, meaning it can never be 100% trusted. If we have to check every single output for something as basic as whether the data we already had is still correct when it comes back to us, then we’re wasting energy and time on mental babysitting instead of using that power for other tasks, including generating the artifact ourselves. (I’ll return to this idea later, in a discussion of prototyping.)

Despite nonstop hype over the last few years, despite C-suite missives to shoehorn AI in wherever it can fit, and despite AI models’ abilities to complete increasingly complex tasks, even current LLMs continue to fail at basic tasks, whether we notice or not.

What causes hallucinations? (LLM history and technology in five minutes)

To critically evaluate LLMs, we should have a basic understanding of how they work. And some history about chatbots is helpful as well. Standing on the shoulders of giants—that is, paraphrasing the work of my colleague Dr. Nitya Mandyam—I’ll tell you about ELIZA, the first chatbot, developed at MIT in the mid-1960s by Dr. Joseph Weizenbaum. 

ELIZA was a rudimentary chatbot built on simple pattern-matching rules, but even so, its conversation partners sometimes experienced an unexpectedly strong emotional connection to the bot. From “ELIZA” on Wikipedia:

Weizenbaum first implemented ELIZA in his own SLIP list-processing language, where, depending upon the initial entries by the user, the illusion of human intelligence could appear, or be dispelled through several interchanges. Some of ELIZA’s responses were so convincing that Weizenbaum and several others have anecdotes of users becoming emotionally attached to the program, occasionally forgetting that they were conversing with a computer. Weizenbaum’s own secretary reportedly asked Weizenbaum to leave the room so that she and ELIZA could have a real conversation. Weizenbaum was surprised by this, later writing: “I had not realized… that extremely short exposures to a relatively simple computer program could induce powerful delusional thinking in quite normal people.”

The key idea that ELIZA demonstrates is that humans are naturally drawn to language-based robots. It’s very easy for us to humanize them, even when they are rudimentary. 

Today’s chatbots have evolved from rule-based systems like ELIZA to count-based models, and then to semantic ones. The foundation of these language models is natural language processing (NLP), which allows us to quantify language into usefully sized bits of information and then do math with those bits. 

One of the breakthroughs in the math of semantic, and later contextual, models is that they use vectors in abstract space, which allows us to map semantic meaning to words. To a count-based language model, “lead” is a token that can co-occur with “pipe” or “president,” among other words. To a contextual model, “lead” as in “pipe” is a completely different vector from “lead” as in “president,” because the words are quantified based on their contexts, not their characters. 

A key aspect of this difference is that it moves us from deterministic models, where outputs can only be sequences that already existed in the training data (anything else has a probability of zero), to probabilistic models, where outputs that are absent from the training data can occur because they have a nonzero probability. 

We can generate text that wasn’t in the training data. This is why it’s called “generative AI.” While this breakthrough in semantic language modeling is extremely powerful, it’s also the Achilles’ heel of the model. We always have the potential to generate text that is not in the training data, which might also mean that it’s untrue or impossible. 

This helps to explain both how the models are so powerful and can write such durable, readable text outputs, and how they are so unreliable and continue to hallucinate. There’s way more to the technical side, but the core idea here is that an LLM’s generative capability is the inexorable partner of its hallucination problem.

Part 2 of the problem: lack of transparency

Remember that Claude example? The issue of how Claude presented its work is another core problem. It shows us that the perception of a completed task is a more important output than an accurately completed task. 

LLMs are not just technologies. They are also products. LLM products such as ChatGPT or Claude are owned by companies—OpenAI and Anthropic, respectively. Their first goal is to keep you using their platform and convert you to a paid customer. Any other stated goal will always be secondary.

We know, from studying ELIZA, that humans are naturally drawn to become attached and form emotional connections with language-based robots. We find them to be really compelling, and we are very willing to believe them, identify with them, and trust them. This, in addition to their actual use cases, makes them a very sellable product. 

And the companies need to sell it. AI spending has far outpaced AI revenue in four successive years, and the gap between the two is growing. I asked Claude to make this graph comparing AI spending vs. revenue for OpenAI and Anthropic from 2022 to 2025, and felt just a bit of schadenfreude:

Two bar charts titled "AI spending vs. revenue, 2022–2025," one for OpenAI and one for Anthropic. A subtitle notes the figures are annual, approximate, and based on reported leaks and analyst estimates. In both, revenue appears in blue and total spending in orange-red. For OpenAI, on a scale running to $25B, both bars sit under $1B in 2022; in 2023, revenue is about $1.5B against $4B in spending; in 2024, about $3.5B against $9B; in 2025, about $13B against $22B. A note reports that OpenAI spends roughly $1.60 to $1.70 for every $1 it earns and expects to remain unprofitable through at least 2028. For Anthropic, on a scale running to $10B, revenue is negligible through 2023 while spending reaches about $1B; in 2024, revenue is about $1B against $3B in spending; in 2025, about $5B against $10B. A note reports spending of roughly two to three times revenue and an expected break-even around 2027–2028.AI spending vs. revenue at OpenAI and Anthropic, 2022–2025. Source: reporting by CNBC, The New York Times, The Wall Street Journal, and Bloomberg. Charts generated by Claude, prompted by Eva Sibinga.

The companies are highly incentivized to sell and push their product. They need to capture market share and convert users to paid customers during this period of growth and relative investor confidence. This matters because the models are designed to create a feeling of convenience. Admitting that they have failed or cannot accomplish a task is not good for their bottom line and runs counter to their pie-in-the-sky marketing, so they don’t do it. 

LLMs are made by companies that value your perception of task completion more highly than they value the accuracy of task completion. The number-one goal is “task appears to be completed,” not “task is completed correctly.” This is why Claude uses the exact same cheerful tone when it returns garbage as when it returns something actually super useful.

That’s not even to mention its baseline complimentary, ingratiating tone. This sycophancy really, truly makes it harder to interpret LLM results correctly. We are not evolutionarily prepared for this type of social interaction, because if a human did this to us consistently, there would be social consequences. You wouldn’t keep going back to a colleague for help if they did the work wrong, cheerfully lied about it, and then gaslit you.

This lack of transparency is antithetical to how we do good data work.

Actual generative intelligence: collective knowledge created in the workshop

I won’t include a lot from the interactive part of the workshop, which was a group ideation session in which we spent about an hour answering questions individually and discussing them together. The goal was for us to reflect on the experiences we had already had with AI, think about aspects of our jobs that we did and did not want automated, ask technical and other questions, and explore the emotions and assumptions that we bring to AI conversations. 

When I opened up a space to be critical or skeptical, there was so much more of this energy than I anticipated. In this room, most of it came from people who had resisted the technology (some continued to; others had recently tried it out), but skepticism or negative feelings also came from people who used the technology daily and relied on it to get through a heavy workload. Some problems were reframed as issues with the workplace (devaluing of labor and expertise, understaffing and underfunding), not inherent issues with how AI technology works.

We recognized that there are successful AI use cases, especially in code-based tasks where the LLM output is code that can be better tested for accuracy and efficiency, and where the developer knows how to think programmatically and evaluate code output. (Although this, too, is rife with problems.) However, the reduction in nuance and idea building is a consistent blocker to implementations that meet our standards for non-AI work in data visualization.

So many new-to-me ideas came up—about the gender pay gap; about wondering if feeling proud of work an AI tool did was “allowed”; about the depths of automation fears from within a group of exceedingly competent business owners, analysts, and consultants. The workshop group happened to be a room full of women, and a feminist thread emerged in the conversation: questioning some of the ways we had seen people on LinkedIn (often but not always men) speak about AI, and talking about which ways did or did not resonate. I mean feminist in the Donna Haraway sense: that seeing everything from nowhere is a f—ed-up god trick, that all knowledge is situated, that disembodied “facts” can lead us further from useful truths (“Situated Knowledges,” 1988).

The common AI stance “get on board or get left behind” sat poorly with the group as a whole. It often lacks nuance and is blind to the skills that create exceptional work in data analysis and visualization. It’s not to say that AI tools can’t be helpful in data visualization. It’s that chasing them from a place of fear and inadequacy feels not only terrible but also incorrect.

We were collectively buoyed by this “AI therapy session,” as one participant called it. Space to openly acknowledge the significant failures of AI tools not only validated lots of experiences in the room, but it also made the genuine use cases for AI tools feel more rewarding to explore.

One more thought on what to do about it

So now we arrive at the challenges—hallucination, intentional lack of transparency—armed with more information. 

If we choose to (or must) use the technology for data visualization work, we can mitigate some of this harm with pointed questions. On accuracy: Is it true and verified? On security: Is my data safe and private? On sovereignty: Am I using the tool the way I want to?

I’ll leave you with one more solution and something else to read. Another solution we talked about in this workshop was the importance of prototyping. I cited Frank Elavsky’s excellent blog post from earlier this year, “On genAI: Was prototyping really a bottleneck?” in which he explores “what if the slow parts about prototyping are actually what make it worth doing?”

The key point here is that prototyping is where we test the intellectual rigor of an idea. I loved this graphic he included about the intellectual refinement and technical refinement of ideas. It makes it clear how incorporating AI early on in the idea stage can make shoddy ideas look sleek while really generating slop, and how the better opportunity for incorporating AI is the “zone of missing skills + resources” for ideas that have already proved themselves in a prototyping stage.

Diagram titled "The prototype slop-zone." The horizontal axis is how technically refined an artifact is; the vertical axis is how intellectually refined the idea is. Low-fidelity prototypes occupy the lower left, mid-fidelity the middle, and high-fidelity the upper right, with a small slice at the far upper right labeled "no longer a prototype." An annotation at the upper left, "what users want genAI for, for their brilliant ideas," points to a gray area labeled "zone of missing skills and resources." A large red region across the lower right is labeled "warning: slop zone — stuff you don't understand but looks pretty good," and an annotation points to it reading "what genAI enables."“The prototype slop-zone,” from “On genAI: Was prototyping really a bottleneck?” Illustration by Frank Elavsky, used with permission.

It’s hard to refine ideas with an AI tool. You have to be in the driver’s seat, because the tool is a reflection of what you ask for paired with a repackaged, statistically likely output of what others have already said on the topic. Elavsky puts it so well: “[P]eople tend to assume that the ideas they have in their head are really good, if they aren’t used to rigorously iterating on ideas.”

A partner who is incentivized to keep us on their platform isn’t one who will say “I just don’t like this direction” to a scrappy drawing in a notebook. Or one who will say “Oh wow, I’ve never thought about it this way before, but you’ve hit on something that’s pretty key.” But both of those kinds of feedback are actually helpful. The LLM partner will be whatever we tell it to be, but always in service of its corporate overlords. 

As Elavsky helps us to see, being thoughtful about when in the process we use AI helps us to develop robust ideas that are worthy of robust technical ends. 

And as this article attempts to lay out, questioning how we engage AI helps us introduce it critically into data visualization work, which demands high standards from visual, numeric, textual, and coding perspectives. 

The best data visualizations prove time and time again that good data analysis is human-centered, asks lots of questions, and makes genuine meaning out of numbers by transforming them into something we can see and contextualize. This part might be sped up or expanded by AI, but it cannot be automated. 

And hooray for that! This is the good part! This is the part that has drawn so many creative, critical, analytical thinkers to data visualization in the first place. Let’s use that same rigorous, flexible thinking when we engage with AI tools. Because if and how we choose to engage with AI in our data visualization workflows has implications for the tools themselves and for our outputs, certainly; but more importantly, if and how we engage with AI reflects how we honor our craft, our own brains, and our lives.

Headshot of Eva

Eva Sibinga

Eva is a freelance web developer, designer, and writer. She’s curious about basically everything, but especially the way diverging experiences lead people to different perceptions of the world. With a background in English and visual art, and an M.S. in Data Analysis & Visualization from The Graduate Center at CUNY, Eva brings a multimedia humanist lens to data-driven questions.

Read the whole story
strugk
35 days ago
reply
Cambridge, London, Warsaw, Gdynia
Share this story
Delete

Dress made of living mycelium can renew and repair itself

1 Share
Dress made from living mycelium textiles from the Shenzhen Institute of Advanced Technology

Researchers in China have created a textile from living mycelium that is self-cleaning, near self-repairing and can be coloured or made UV protective through "plug-and-play" add-ons of different fungi or yeast.

The breakthrough from the researchers at the Shenzhen Institutes of Advanced Technology is a type of engineered living material (ELM) – a material built off living organisms that stay active even after they're fabricated into their final form.

This distinguishes it from most contemporary uses of mycelium, which involve drying and effectively killing the fungus to produce a stable, non-growing material that has become a popular emerging alternative to plastic packaging and vinyl.

Photo of a dark blue dress with white ruffles on the bottom hem, sitting on a mannequinA dress has been made by Peelsphere using material from Ke Li and her colleagues

By working with living but dormant cordyceps militaris fungus instead, the researchers have been able to take advantage of its biological functions. The result is a material that is self-renewing and responsive to its environment, in ways that could one day transform architecture and clothing – as seen in a prototype dress created together with material innovation company Peelshere.

It can also be adapted by mixing in other fungi or yeast, lead researcher Ke Li and her team detail in a paper in the peer-reviewed journal Science Advances.

In it, they describe a "programmable fungal platform" where mycelium is treated like a modular system, with the sheet material forming a base structure and extra biological abilities, such as colour and UV resistance, becoming "plug-and-play" add-ons via other organisms.

Photo of a hand holding a sheet of translucent, caramel-coloured leathery material that is in fact a fungal textileThe fungal textile is made in sheets

This gets their textile closer to the self-repair, environmental responsiveness and controllable functionality that is the promise of engineered living materials, they argue.

"While synthetic biology has greatly expanded the functional capabilities of ELMs, a persistent challenge lies in integrating autonomous structural assembly with sustained biological activity at macroscopic scales," the scientists write.

"Achieving such integration is essential for practical applications, from adaptive textiles to architectural biomaterials, where mechanical robustness, spatial uniformity and scalable fabrication must converge with engineered biological function."

The ELM's self-renewing and semi-repairing functionality comes from the mycelium base structure. Following drying at 45 degrees, the material is not quite living and not quite dead, but instead in a "low-metabolic, dormant-like state", Li told Dezeen, meaning it is not actively growing.

However, new growth can be triggered by applying a nutrient solution of potato water, leading the dormant mycelium to germinate, send out new fungal filaments and renew the material's surface.

When this nutrient solution is applied over a hole, along with a small patch of fresh fungus, it triggers the living cells to grow across the gap, seamlessly repairing the surface without any adhesives or stitching. The material is also naturally self-cleaning, as it is hydrophobic.

Photo of an ornamental butterfly made of wire with a translucent deep blue textile filling in its wingsKe Li also demonstrated the material in self-pigmented blue on the wings of a butterfly ornament

The blue colour and UV resistance, meanwhile, come from brewer's yeast – also known as saccharomyces cerevisiae – and aspergillus niger fungus, respectively. These are mixed in with the cordyceps militaris at the beginning and simply co-cultured, avoiding the need for genetic engineering.

A prototype dress has been made out of the scientists' material by Berlin-based Peelshere, whose founder YouYang Song is a friend of Li's. Song developed the conceptual and aesthetic design of the dress, while her China-based colleague Ruochen Wang took care of the cutting and construction.

The dress features several versions of the material, including some co-cultured with brewer's yeast for the consistently self-pigmented blue shade in the body of the garment.

Photo of a square plant pot made of earthy brown leather-like material, holding a succulentThe material can also be shaped into structures like this box

Li told Dezeen that the material is suitable for applications such as conceptual fashion, accessories, decorative textile surfaces, exhibition pieces and biodegradable packaging.

"Its distinctive surface texture, biological colouring, controlled repair and biodegradability may be particularly useful in applications where visual expression and a defined product lifetime are important," she said.

"Further improvements in durability, moisture resistance, safety and manufacturing consistency would be needed before it could be considered for routine clothing or permanent architectural use."

Peelsphere's main product is a plant-based and waterproof leather alternative made of fruit peels and algae.

Photography by Ke Li.

Read the whole story
strugk
43 days ago
reply
Cambridge, London, Warsaw, Gdynia
Share this story
Delete

Ailments – Potential AI-Induced Mental & Behavioural Disorders — Information is Beautiful

1 Share

Generative AI isn’t just changing how we work. It’s may also be affecting how we think, create, collaborate, procrastinate, and occasionally lose the plot.

You may have heard of AI Mania and AI Psychosis. 

Here’s a field guide to more potential AI-induced mental habits, compulsions and behavioural quirks that may be emerging in the age of AI.

None are recognised medical conditions.

Many might feel uncomfortably familiar.

Did we miss any? Suggest one

Written and designed by David McCandless. Additional contributions from Nik Roope, Cal Newport, Hanna Piotrowska


Some examples of AI induced mental habits & compulsions

AI Psychosis
Total, unquestioned ‘buy in’ to the idea that AI is a must-have super-human replacement for anything human: creativity, cognition, coding, customer-relations, companionship, conversation – and common sense. 

NarcAIssism
Euphoric ego-inflation induced by prolonged interaction with overly agreeable AI chatbots. The sufferer develops an outsized certainty of their creative brilliance, personal insight and strategic acuity.

PolyLLMory
The inability to commit to a single AI model, leading to the insertion of all prompts into Claude, ChatGPT & Gemini simultaneously. 

Though there may be a ‘primary’, relationships with all three models remain technically ‘open’. 

FOFAB
Fear of Falling Behind. Ambient dread and background anxiety over the possibility that everyone else and their hairdresser are using AI more effectively than you.  And of becoming professionally obsolete by Tuesday.

Agentic Burgerflipping
Spending entire workdays supervising AI bots & agents rather than doing any actual skilled work. 

Today’s tasks: 1) evaluating outputs 2) re-prompting after errors and 3) typing “fix it” followed by the return key.

Claudependency
Irrespective of the task or issue – personal, professional, large, small – a <strong>Claudependent</strong> must ‘discuss’ it first with their Anthropic LLM, reporting back their decision with “My AI said…”

UpSkill Sisyphus
Constant efforts to learn different AI models, test new tools, and stay current generates a permanent cognitive churn. Often laundered as “adaptability” but actually an exhausting, perma-treadmill with no destination.

Cognitive Laxity
Atrophy of memory, reasoning and problem-solving muscles due to habitual outsourcing of mental effort to AI. 

Example: all universities graduates from 2025 onwards

AI Burnout
Exhaustion caused by the endless steering, correcting and re-prompting of mediocre AI output. 

Did we miss any? Suggest one

Read the whole story
strugk
55 days ago
reply
Cambridge, London, Warsaw, Gdynia
Share this story
Delete

Iowa Hog Barn Becomes Mushroom Farm as Transfarmation Opens Second Demonstration Hub - vegconomist - the vegan business magazine

1 Share

A former hog operation in Radcliffe, Iowa, has been converted into a specialty mushroom farm and now serves as the second demonstration hub for The Transfarmation Project, the farm-transition program run by nonprofit Mercy For Animals.

The site, known as 1100 Farm, was previously a concentrated animal feeding operation (CAFO) that raised an estimated 8,000 pigs per year. Working with the Transfarmation team, contractors, and consultants, the Faaborg family converted one of the property’s hog barns into a growing space for specialty mushrooms. Transfarmation defines a demonstration hub as a former CAFO that has been repurposed into a specialty crop farm, used for research, farmer visits, and documenting the economics of transitioning away from industrial animal agriculture.

Three decades of hog farming before the switch

Tammy and Rand Faaborg raised pigs on the land for 30 years, starting with a small number of animals before building two barns under an integrator contract, each housing 1,100 pigs. The farm’s name references that figure. The couple began working with Transfarmation in June 2021 and, a year later, decided to stop raising pigs and pursue a full conversion. Transfarmation awarded the family a grant in 2023 to pilot a mushroom-growing project.

“It takes unimaginable courage to look at a multi-decade family business and say, ‘We need to find a better way.'”

The family expanded from fresh mushrooms into mushroom tinctures, jerky, coffee, and hot chocolate blends, and began producing their own mushroom blocks. A 2025 feature in The New York Times drove close to 1,000 orders in the days following publication, according to Transfarmation. The barn overhaul followed, with the Faaborgs handling most of the construction themselves.

Transfarmation© Transfarmation

Iowa’s factory farming footprint

The location carries weight for the organization. Iowa leads US production of pork, poultry, eggs, and corn, and is home to over 5,000 pig farms. Hogs outnumber people in the state by about seven to one. Transfarmation states that Iowa’s factory farms generate around 300 million pounds of manure daily, roughly 25 times the waste produced by the state’s human population, and links the resulting nitrate levels in waterways to the state’s agricultural output.

Transfarmation was founded in 2019 by Mercy For Animals president Leah Garcés and has supported farm transitions across states including Indiana, Iowa, North Carolina, and Texas, with most participating farmers moving into specialty mushrooms. The 1100 Farm site follows the program’s first demonstration hub, a converted poultry farm in North Carolina.

Katherine Jernigan, Director of Transfarmation, stated, “It takes unimaginable courage to look at a multi-decade family business and say, ‘We need to find a better way.’ But that is exactly what the Faaborgs did. This hub is a living testament that we don’t have to tear down our agricultural heritage to build a better future. The Faaborgs have shown us that a food system that works for farmers, animals, and the planet isn’t just a dream, it’s already happening.”

More news from the region:

Newsletter

Subscribe for the vegconomist-newsletter and regularly receive the most important news from the vegan business world.

Subscribe

Read the whole story
strugk
55 days ago
reply
Cambridge, London, Warsaw, Gdynia
Share this story
Delete
Next Page of Stories