<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Partech Systems — Journal</title>
    <link>https://partechsystems.com/blog/</link>
    <description>Partech Systems is an AI consultancy. We help teams work out where AI genuinely earns its keep — and where it doesn&#39;t — then build the thing properly.</description>
    <language>en-GB</language>
    <lastBuildDate>Thu, 03 Sep 2026 12:55:12 +0000</lastBuildDate>
    <atom:link href="https://partechsystems.com/feed.xml" rel="self" type="application/rss+xml"/>
    <item>
      <title>Your review process was built for a world where output was hard</title>
      <link>https://partechsystems.com/blog/performance-reviews-after-ai/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/performance-reviews-after-ai/</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhavna Ate</dc:creator>
      <category>Performance</category>
      <category>People</category>
      <category>AI</category>
      <category>Management</category>
      <description>When everyone&#39;s output doubles, output stops telling you anything. Most performance systems haven&#39;t noticed yet, and the ones quietly adjusting are adjusting in the worst possible way.</description>
      <content:encoded><![CDATA[&lt;p&gt;A head of engineering said something to me in the spring that I&#39;ve been chewing on ever since. He was looking at his team&#39;s half-year numbers and he said, almost apologetically, &#34;everyone&#39;s a top performer now and I don&#39;t believe it.&#34;&lt;/p&gt;
&lt;p&gt;He wasn&#39;t being unkind. The numbers had genuinely gone up — throughput, tickets, merged work, all of it — across the whole team, including the people he&#39;d have told you six months earlier were struggling. His review process was going to hand out uniformly strong ratings, his budget wouldn&#39;t cover uniformly strong outcomes, and he had no defensible basis on which to differentiate.&lt;/p&gt;
&lt;p&gt;That&#39;s the problem in miniature, and it is arriving in a lot of organisations at once. Most performance systems, whatever they claim in the framework document, ultimately reward visible output. That was a reasonable proxy for as long as producing output was the expensive part. It isn&#39;t any more, and a proxy that everyone can now satisfy has stopped being a measurement.&lt;/p&gt;
&lt;h2 id=&#34;the-two-bad-reflexes&#34;&gt;The two bad reflexes&lt;/h2&gt;
&lt;p&gt;I&#39;ve seen two instinctive responses and I&#39;d argue against both.&lt;/p&gt;
&lt;p&gt;The first is to raise the bar quietly. If everyone&#39;s doing twice as much, expect twice as much, and carry on with the same scale. This is what happens by default, because it requires no decision and no announcement. It&#39;s also how you get a workforce that received a genuinely useful tool and experienced it as an increase in workload with no acknowledgement — which is a fast route to the sort of resentment that doesn&#39;t show up in an engagement survey until it&#39;s structural. If you are going to reset expectations, reset them once, deliberately, out loud, with the reasoning attached. Drift is the cruellest way to do it.&lt;/p&gt;
&lt;p&gt;The second reflex is to start measuring AI use itself. Put &#34;adoption of AI tools&#34; in the objectives, track licence usage, reward the enthusiastic. I understand the impulse — leadership wants to see the investment landing — but you are measuring an input, and the moment you attach a rating to an input you get theatre. People will use the tool for things it isn&#39;t good for, and the person who correctly judged that their work didn&#39;t need it gets marked down for good judgement. I&#39;ve already watched this happen at one organisation. Six months of impressive dashboards and no discernible change in anything that mattered.&lt;/p&gt;
&lt;h2 id=&#34;whats-actually-left-to-distinguish-people-by&#34;&gt;What&#39;s actually left to distinguish people by&lt;/h2&gt;
&lt;p&gt;If output is now cheap and roughly equal, what&#39;s expensive and unequal?&lt;/p&gt;
&lt;p&gt;Judgement, mostly. Which is a soft word for a set of quite hard, specific things.&lt;/p&gt;
&lt;p&gt;Choosing the right problem. When it takes an afternoon rather than a fortnight to produce a thing, the cost of producing the wrong thing drops too — so more wrong things get produced, and the person who works out what&#39;s actually worth doing becomes disproportionately valuable. That has always been true. It&#39;s just no longer masked by the effort involved in execution.&lt;/p&gt;
&lt;p&gt;Knowing when the answer is wrong. This is the one I&#39;d weight most heavily, and it&#39;s the hardest to see. The engineer who noticed that the generated code handled the timezone case incorrectly, the analyst who spotted that the summarised figures had double-counted a region — they prevented something, and prevention is invisible in every performance system I&#39;ve ever read. If you don&#39;t deliberately go looking for it, you will rate the person who shipped fast above the person who stopped them shipping wrong.&lt;/p&gt;
&lt;p&gt;Deciding not to. Restraint has never been rewarded in performance reviews and it&#39;s more valuable now than it&#39;s ever been.&lt;/p&gt;
&lt;h2 id=&#34;how-to-evidence-that-without-it-becoming-a-personality-contest&#34;&gt;How to evidence that without it becoming a personality contest&lt;/h2&gt;
&lt;p&gt;Here&#39;s my real worry, and I want to name it plainly rather than let it sit under the surface.&lt;/p&gt;
&lt;p&gt;When you shift assessment from countable output to judgement, you shift it towards things that are argued for rather than demonstrated. And the people who are good at arguing for their own contribution are not a random sample. They skew towards the confident, the fluent, the socially central, the ones who look like the person doing the assessing. Every organisation I&#39;ve worked in has had at least one quiet, excellent person whose value was obvious to their team and invisible to the process. Making the criteria more subjective makes that worse, systematically, and it makes it worse for exactly the groups your inclusion data is already telling you about.&lt;/p&gt;
&lt;p&gt;So the criteria have to get more subjective and the evidence has to get more concrete, at the same time. That&#39;s the needle.&lt;/p&gt;
&lt;p&gt;Practically, what I&#39;d change. Ask for decisions, not volume — three things you decided this half, what the alternatives were, and how it turned out. Ask managers to record catches as they happen, in a line or two, in the moment, because nobody reconstructs them at review time. Ask peers a specific question rather than a general one: not &#34;rate this person&#39;s collaboration,&#34; but &#34;when did this person save you from something?&#34; And separate the two questions that reviews currently conflate — did the organisation get value from this area of work, which is a team-level question, and did this individual work well, which is not the same thing and increasingly isn&#39;t even correlated.&lt;/p&gt;
&lt;p&gt;One more, and it applies to anyone with a promotion ladder. Most ladders assume that doing the junior work well is how you demonstrate readiness for the senior work. If the junior work is now largely automated, that evidence path has quietly closed, and you will find yourself with a cohort of people who have no way to prove they&#39;re ready. That&#39;s a design problem in the ladder, not a shortcoming in the people, and it needs fixing before the first promotion round where it bites.&lt;/p&gt;
&lt;p&gt;None of this is a small edit to the form. It&#39;s a change to what you believe good work is, which is why most organisations will do the quiet-bar-raise instead. I&#39;d just note that the ones who take it seriously will be able to answer the question my engineering friend couldn&#39;t, and the ones who don&#39;t will spend the next few years handing out ratings they don&#39;t believe.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The customer is the daughter, not the mother</title>
      <link>https://partechsystems.com/blog/ai-ageing-and-care/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/ai-ageing-and-care/</guid>
      <pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Dimple Paratey</dc:creator>
      <category>Ageing</category>
      <category>Care</category>
      <category>AI</category>
      <category>People</category>
      <description>Almost every AI product sold for older people is really sold to their adult children. Once you notice that, you can&#39;t unnotice it — and it explains why so much of it feels wrong to the person it&#39;s aimed at.</description>
      <content:encoded><![CDATA[&lt;p&gt;My aunt is eighty-three and lives on her own in a terraced house in Leicester that she has no intention of leaving. There is a family WhatsApp group, and a good proportion of what happens in it is her children and nieces quietly worrying about her while being careful not to let her see us doing it.&lt;/p&gt;
&lt;p&gt;Last year one of my cousins bought her a sensor system. Little discs on the doors and the kettle and the medicine cupboard, an app on our phones, and a gentle promise that we&#39;d know if something was wrong. It worked exactly as advertised. We could see when she&#39;d got up, when she&#39;d made tea, whether she&#39;d opened the pill drawer.&lt;/p&gt;
&lt;p&gt;She hated it. Not loudly — she&#39;s not a loud person — but within about six weeks the discs had been &#34;knocked off by the cleaner,&#34; all of them, in an impressive coincidence.&lt;/p&gt;
&lt;p&gt;I&#39;ve thought about that a great deal since, because it taught me something I now see everywhere in this market. The product worked. It just wasn&#39;t for her.&lt;/p&gt;
&lt;h2 id=&#34;whos-actually-paying&#34;&gt;Who&#39;s actually paying&lt;/h2&gt;
&lt;p&gt;Almost every piece of AI-adjacent technology sold into ageing is bought by the adult child and used by the parent. The buyer&#39;s problem is anxiety. The user&#39;s problem is entirely different — she isn&#39;t anxious, she&#39;s fine, and what she wants is for everyone to stop treating her like a project.&lt;/p&gt;
&lt;p&gt;Once you see this you can&#39;t unsee it, and it explains why so much of the category feels slightly off. The features are optimised for the reassurance of the person holding the phone: dashboards, streaks, alerts, activity summaries. The person in the house gets no interface at all. She&#39;s not a user of that product. She&#39;s the data source.&lt;/p&gt;
&lt;p&gt;Which is a very strange thing to do to an adult who ran a school office for thirty years and does the cryptic crossword faster than I do.&lt;/p&gt;
&lt;p&gt;I&#39;m not against remote monitoring — falls are genuinely how independence ends for a lot of people, and a system that gets someone found in twenty minutes instead of nine hours is not a gimmick. But there&#39;s an enormous difference between &#34;this alerts someone if I fall&#34; and &#34;this reports my day to my children,&#34; and the products blur it deliberately, because the blur is what sells.&lt;/p&gt;
&lt;h2 id=&#34;what-she-does-use&#34;&gt;What she does use&lt;/h2&gt;
&lt;p&gt;Here&#39;s the interesting part. She&#39;s not a technophobe. She uses things constantly — they&#39;re just the ones where she&#39;s the one being helped.&lt;/p&gt;
&lt;p&gt;She photographs letters and has them read out to her, because the print on the pension correspondence is absurd and her eyes are tired by four o&#39;clock. She dictates messages rather than typing them, at length, with punctuation spoken aloud in a way I find delightful. After a hospital appointment last winter she recorded the consultation, with permission, and got a plain summary of it afterwards — which mattered because she&#39;d taken in about a third of what the registrar said, as most people do in that chair.&lt;/p&gt;
&lt;p&gt;None of that is marketed to older people. There&#39;s no &#34;silver&#34; branding on any of it. It&#39;s just ordinary technology that happens to remove a specific, concrete obstacle, and she found most of it herself.&lt;/p&gt;
&lt;p&gt;The pattern seems to be: things that restore capability get adopted, things that transfer oversight to someone else get resisted. That&#39;s not a subtle distinction and yet the sector keeps building the second kind.&lt;/p&gt;
&lt;h2 id=&#34;the-companion-question&#34;&gt;The companion question&lt;/h2&gt;
&lt;p&gt;I&#39;ll be honest that I don&#39;t know what I think about the companionship products, and I&#39;d distrust anyone who&#39;s certain.&lt;/p&gt;
&lt;p&gt;Loneliness in older people is a genuine health crisis, not a soft issue — the physiological effects are on par with things we take extremely seriously. A device that talks to someone every day, remembers what they said yesterday, asks about the grandchildren by name, is measurably better than the silence it replaces. I&#39;ve read the research and I&#39;ve spoken to people who work in care homes, and they&#39;re mostly not sentimental about it. If it helps, it helps.&lt;/p&gt;
&lt;p&gt;What I can&#39;t get past is the direction it lets things drift. It&#39;s very easy to imagine a commissioning decision where a companion product is funded because it&#39;s cheaper than the visiting service it quietly displaces, and nobody ever writes down that trade because nobody ever proposes it out loud. That&#39;s how it would happen. Not a decision — a substitution, made on a spreadsheet.&lt;/p&gt;
&lt;p&gt;So my position, for now, is: fine as an addition, never as a saving. And I&#39;d want the person it&#39;s for to have been asked, in a real conversation, whether they want it.&lt;/p&gt;
&lt;h2 id=&#34;the-test-id-apply&#34;&gt;The test I&#39;d apply&lt;/h2&gt;
&lt;p&gt;If you&#39;re building in this space, one question. Can she turn it off?&lt;/p&gt;
&lt;p&gt;Not &#34;is there a setting.&#34; Can the eighty-three-year-old woman in the house decide, without a phone call to her family, without a fuss, without anyone being informed, that today she&#39;d rather not be observed — and does the product still work for her afterwards?&lt;/p&gt;
&lt;p&gt;If yes, you&#39;ve built something for her. If no, you&#39;ve built something for us, and you should at least be honest in the marketing about which one it is.&lt;/p&gt;
&lt;p&gt;My aunt is fine, by the way. She&#39;s had a wrist alarm since the spring, which she chose herself, in the shop, having asked several sharp questions about who gets called first. That one&#39;s still on.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>It was fine in March</title>
      <link>https://partechsystems.com/blog/it-was-fine-in-march/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/it-was-fine-in-march/</guid>
      <pubDate>Thu, 28 May 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>Operations</category>
      <category>Monitoring</category>
      <category>LLMs</category>
      <category>Reliability</category>
      <description>AI systems don&#39;t fail loudly. They get quietly worse while every dashboard stays green, and you find out from a customer. Here&#39;s what to watch instead, and which signal is worth more than all the others.</description>
      <content:encoded><![CDATA[&lt;p&gt;The sentence I hear most often when something has gone wrong with a production AI system is: &#34;It was fine in March.&#34;&lt;/p&gt;
&lt;p&gt;It probably was. That&#39;s the problem. Ordinary software fails in a way that sets off a pager — an exception, a timeout, a queue backing up. A system with a model in it degrades instead. The answers get a bit vaguer, the classifications drift a couple of points, the extraction starts missing one field on documents from one supplier. Every dashboard stays green, because every dashboard is measuring whether the thing responded, and it did. It responded beautifully. It was just increasingly wrong.&lt;/p&gt;
&lt;p&gt;Then in June somebody escalates a complaint, you go and look, and you find the decline started in April.&lt;/p&gt;
&lt;h2 id=&#34;three-ways-it-rots&#34;&gt;Three ways it rots&lt;/h2&gt;
&lt;p&gt;Worth separating these, because they need different responses.&lt;/p&gt;
&lt;p&gt;The inputs change. You built a support classifier on last year&#39;s tickets; since then the company launched two products, the help centre was rewritten, and a chunk of traffic moved to mobile where people type shorter and worse. Nothing about the model changed. The world it was fitted to did.&lt;/p&gt;
&lt;p&gt;The right answer changes. This one&#39;s nastier because the inputs look identical. A fraud pattern evolves specifically to look normal. A policy is updated, so a document that was compliant in January isn&#39;t in May, and the model is confidently applying the old rule with the old confidence.&lt;/p&gt;
&lt;p&gt;And — by a distance the most common in practice — somebody upstream changed something. A form got a new optional field. An API started returning nulls where it used to return empty strings. A supplier switched their invoice template. A PDF library was upgraded and now the tables extract in a different column order. None of this is an AI problem at all, but it presents as one, and teams lose weeks investigating the model when the culprit was a release note nobody read.&lt;/p&gt;
&lt;p&gt;There&#39;s also the reflexive case, which I&#39;d file under &#34;know that it exists.&#34; If your model&#39;s outputs shape the data you later learn from — it ranks what people see, so people click what it ranked — then it&#39;s slowly training on its own opinions. That takes longer to hurt you and is much harder to unpick.&lt;/p&gt;
&lt;h2 id=&#34;what-to-watch&#34;&gt;What to watch&lt;/h2&gt;
&lt;p&gt;Not accuracy, mostly, because you rarely have the ground truth in time. Watch the things that move &lt;em&gt;before&lt;/em&gt; accuracy does.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Input distribution.&lt;/strong&gt; Length, language, source system, document type, the share of requests hitting each category. You don&#39;t need anything statistically elegant here; a weekly comparison against a fixed reference window catches almost everything worth catching. When the shape of what&#39;s arriving changes, something changed upstream, and you want to know that in week one rather than month three.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Output distribution.&lt;/strong&gt; If a classifier that has always put 12% of tickets in &#34;billing&#34; is putting 19% there this week, either your customers have had an unusual week or you have a problem. Both are worth a look.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fallback and refusal rates.&lt;/strong&gt; How often does it decline, hedge, return an empty result, or trip a guardrail? This is a superb early sensor because it moves first and it&#39;s trivial to log.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Human override rate.&lt;/strong&gt; If there&#39;s a person in the loop — an agent who can edit the draft, a reviewer who can reject the extraction — the rate at which they change the machine&#39;s answer is the single most valuable number you have. It&#39;s free continuous evaluation performed by domain experts, and I&#39;m consistently amazed how many teams collect it and never plot it. Plot it. Break it down by segment. It will tell you where the system is failing long before any aggregate metric does.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cost per finished task.&lt;/strong&gt; Not per request. When the model starts struggling, retries rise, escalations rise, and the bill drifts up before quality visibly drops. Cost is a quality signal in disguise.&lt;/p&gt;
&lt;h2 id=&#34;making-the-alarm-real&#34;&gt;Making the alarm real&lt;/h2&gt;
&lt;p&gt;Two things separate teams who catch this from teams who don&#39;t.&lt;/p&gt;
&lt;p&gt;The first is that someone owns it. A named person or rota looks at the numbers weekly, and can stop a rollout. Not a dashboard that exists — dashboards that exist are wallpaper. Somebody whose job includes noticing.&lt;/p&gt;
&lt;p&gt;The second is a small set of held-back cases that get run and graded properly on a schedule, monthly or so, by a human. A hundred examples, drawn from real recent traffic, marked by someone who knows the answer. It&#39;s a couple of hours of work and it&#39;s the only thing that gives you a true accuracy number rather than a proxy. Every team that&#39;s told me they don&#39;t have time for this has later spent considerably longer reconstructing when a regression began.&lt;/p&gt;
&lt;p&gt;Then, when you do change something — new model version, new prompt, new retrieval — run it in shadow first. Both paths execute, only the old one is user-visible, and you compare on live traffic for a week. Practically every unpleasant surprise I&#39;ve seen would have been caught by a week of shadow running, and shadow running costs money rather than reputation, which is the better currency to spend.&lt;/p&gt;
&lt;h2 id=&#34;the-uncomfortable-bit&#34;&gt;The uncomfortable bit&lt;/h2&gt;
&lt;p&gt;Most organisations budget for building an AI system and treat running it as overhead. It isn&#39;t. A system with a model in it needs the same standing attention as any other piece of infrastructure that touches customers, and slightly more, because its failures are polite.&lt;/p&gt;
&lt;p&gt;The good news is that none of the above is hard. It&#39;s a handful of counters, one weekly meeting, a monthly grading session, and the discipline to shadow-run changes. The teams doing it aren&#39;t clever. They&#39;re just the ones who&#39;ve already had the June conversation once and would rather not have it again.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The EU AI Act, a year in: what it actually changed for the people building</title>
      <link>https://partechsystems.com/blog/eu-ai-act-2026/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/eu-ai-act-2026/</guid>
      <pubDate>Thu, 30 Apr 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>EU AI Act</category>
      <category>Compliance</category>
      <category>Governance</category>
      <category>Regulation</category>
      <description>A year in, the EU AI Act has stopped being a legal-team topic and become a build problem. The teams treating it as a compliance review at the end are paying for it twice. Here&#39;s what it actually forces, and the posture that works if you&#39;re not a hyperscaler.</description>
      <content:encoded><![CDATA[&lt;p&gt;The most useful sentence I can give you about the EU AI Act is this: it stopped being a legal-team topic the moment the obligations landed rather than loomed, and it became a build problem. Teams treating it as a compliance document to review at the end are the ones doing expensive rework. Teams treating it as a set of constraints on the build are the ones shipping. I&#39;ve watched both happen this year, and the gap between them is money.&lt;/p&gt;
&lt;p&gt;So let me lay out what the Act actually requires, where the legal reading and the engineering reality split, and a posture that works if you&#39;re a mid-sized company rather than a hyperscaler with a standing regulatory-affairs function.&lt;/p&gt;
&lt;h2 id=&#34;the-risk-tiers-fast&#34;&gt;The risk tiers, fast&lt;/h2&gt;
&lt;p&gt;Four buckets, and almost everything downstream depends on which one you&#39;re in.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Prohibited.&lt;/strong&gt; Social scoring, certain biometric categorisation, manipulative systems. If you&#39;re here, the answer is don&#39;t.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;High-risk.&lt;/strong&gt; Systems used in defined sensitive contexts — recruitment, credit, education, critical infrastructure, certain safety components. This is the tier with teeth, and the one most teams wildly underestimate their exposure to.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Limited-risk.&lt;/strong&gt; Chatbots, generated content. The obligation is mainly transparency: tell people they&#39;re dealing with AI or looking at synthetic media.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Minimal-risk.&lt;/strong&gt; Everything else. No specific obligations, though general law still applies.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The trap is assuming you&#39;re limited-risk because your product &lt;em&gt;feels&lt;/em&gt; low-stakes. The classification follows the use context, not the feeling. A mundane scoring model becomes high-risk the second it&#39;s used in hiring or lending. Map your features to contexts early, because the tier sets the cost of everything else.&lt;/p&gt;
&lt;h2 id=&#34;what-high-risk-actually-forces&#34;&gt;What &#34;high-risk&#34; actually forces&lt;/h2&gt;
&lt;p&gt;This is where the legal summary and the engineering reality are furthest apart. The legal version reads &#34;risk management, data governance, documentation, logging, human oversight, accuracy and robustness.&#34; Translated into work your engineers actually have to do:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Logging that&#39;s auditable, not just present.&lt;/strong&gt; Records of what the system did and on what basis, retained and queryable. If your inference path doesn&#39;t emit structured, retained logs tied to inputs and model versions, that&#39;s a build item, not a config change.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data governance you can evidence.&lt;/strong&gt; Provenance, representativeness, bias testing of training and validation data — documented, not asserted. Most teams can &lt;em&gt;describe&lt;/em&gt; their data. Far fewer can produce the artefact proving they examined it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Real human oversight.&lt;/strong&gt; Not a human who can technically click approve, but one positioned and equipped to actually intervene. A rubber stamp doesn&#39;t satisfy this, and it&#39;s obvious in an audit.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Documentation maintained as the system changes.&lt;/strong&gt; This is the one that rots. Docs written once at launch are wrong within two release cycles.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The theme runs through all four: every one is cheap if it&#39;s in the pipeline and ruinous if you reconstruct it afterwards. You cannot retrofit a year of audit logs you never wrote.&lt;/p&gt;
&lt;h2 id=&#34;general-purpose-models-and-the-transparency-layer&#34;&gt;General-purpose models and the transparency layer&lt;/h2&gt;
&lt;p&gt;If you build on general-purpose models — most of us do — obligations sit both upstream and on you. Providers of the large models carry their own documentation and transparency duties. But you don&#39;t get to outsource your end. You still disclose AI interaction and synthetic content where the limited-risk rules apply, and if your use sits in a high-risk context, the provider&#39;s compliance does not discharge yours. Get clear contractually on what the model provider hands you — model cards, evaluation data, the documentation you can actually rely on — because the gaps become your problem at audit time.&lt;/p&gt;
&lt;h2 id=&#34;the-posture-for-a-mid-sized-company&#34;&gt;The posture for a mid-sized company&lt;/h2&gt;
&lt;p&gt;You don&#39;t have a hyperscaler&#39;s headcount. So buy nothing you can build into the pipeline, and build the cheap controls now instead of the expensive ones later.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Classify first, once, properly.&lt;/strong&gt; A half-day mapping features to tiers saves months. Most of your surface is probably minimal or limited-risk; concentrate the effort on the genuinely high-risk slice.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make logging and versioning a platform capability, not a per-feature scramble.&lt;/strong&gt; If model version, inputs, outputs and the oversight decision are captured by default, compliance for new features becomes near-free. Highest-leverage thing you can do.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Treat documentation as generated, not authored.&lt;/strong&gt; Pull model cards, eval results and data lineage from the systems that already produce them. Hand-maintained docs drift. Derived docs don&#39;t.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Put one accountable owner on it.&lt;/strong&gt; Not a committee. Someone who owns the classification and the evidence and updates both as the product changes.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;governance-is-a-product-problem&#34;&gt;Governance is a product problem&lt;/h2&gt;
&lt;p&gt;Here&#39;s the framing I keep coming back to. People file the AI Act under legal risk, alongside GDPR or a SOC 2. Category error. The obligations — logging, oversight, data governance, documentation — are all properties of how the system is built and run. They live in the pipeline. Which makes them an engineering responsibility that legal verifies, not a legal responsibility that engineering occasionally helps with.&lt;/p&gt;
&lt;p&gt;Bake it in and the marginal cost per feature trends towards zero. Bolt it on and you pay twice: once to build the feature, again to reconstruct the evidence it was built responsibly. The Act didn&#39;t create that choice. It just made the bill for the wrong one payable.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>If AI does the junior work, where do your seniors come from?</title>
      <link>https://partechsystems.com/blog/ai-and-junior-roles/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/ai-and-junior-roles/</guid>
      <pubDate>Mon, 20 Apr 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhavna Ate</dc:creator>
      <category>Careers</category>
      <category>People</category>
      <category>AI</category>
      <category>Talent</category>
      <description>AI is brilliant at exactly the work junior people used to cut their teeth on. Automate all of it and you save money now while dismantling the pipeline that produces your future seniors. You can&#39;t hire your way out of that, because seniors are grown.</description>
      <content:encoded><![CDATA[&lt;p&gt;Here&#39;s a trap that looks like a win on a spreadsheet. AI is genuinely good at the work we used to hand to junior people — the first-draft code, the research summary, the basic analysis, the routine document. So the obvious move, the one finance loves, is to use the tool instead of hiring the junior. Same output, lower cost, no onboarding. The maths is clean.&lt;/p&gt;
&lt;p&gt;The maths is also lying to you, because it only counts this year.&lt;/p&gt;
&lt;h2 id=&#34;seniors-are-grown-not-procured&#34;&gt;Seniors are grown, not procured&lt;/h2&gt;
&lt;p&gt;You cannot buy a senior. You can buy &lt;em&gt;a&lt;/em&gt; senior — poach one from someone else — but as an industry, as an economy, the supply of senior people exists only because somewhere a junior person spent years doing unglamorous work badly, then less badly, then well, and became the senior. There is no other path. Expertise is the residue of having done the work, including the boring and the wrong parts of it.&lt;/p&gt;
&lt;p&gt;The junior tasks we&#39;re so keen to automate aren&#39;t just output. They&#39;re the apprenticeship. The analyst grinding through a tedious dataset is learning what clean data feels like and developing the instinct that something&#39;s off before they can say why. The junior developer fixing trivial bugs is building a map of the codebase in their head. That &#34;low-value&#34; work is where judgement is manufactured. You&#39;re not paying for the output. You&#39;re paying for what doing the output does to the person.&lt;/p&gt;
&lt;p&gt;Take it away entirely and you get a cohort who can prompt a model to produce a passable first draft but have never built the underlying judgement to know when the draft is subtly, dangerously wrong. They skipped the years that grow the instinct. And nobody notices the gap for about five years — right up until your current seniors retire and there&#39;s nobody behind them, because you stopped making any.&lt;/p&gt;
&lt;h2 id=&#34;the-pipeline-problem-is-everyones-and-no-ones&#34;&gt;The pipeline problem is everyone&#39;s, and no one&#39;s&lt;/h2&gt;
&lt;p&gt;This is what economists call a collective action problem, and it&#39;s worth naming because it explains why it&#39;ll happen by default. For any single company, cutting junior hiring is rational — let everyone else bear the cost of training people, then hire the finished article. But if everybody reasons that way, nobody trains anyone, and in a decade there&#39;s a shortage of mid-level and senior talent that no salary can conjure, because the people simply don&#39;t exist. Each firm optimised itself into a problem none of them can buy their way out of.&lt;/p&gt;
&lt;p&gt;I&#39;m not appealing to anyone&#39;s better nature here. I&#39;m pointing out that the firm which keeps growing its own people, while competitors strip-mine the pipeline, ends up with the one thing that becomes genuinely scarce. That&#39;s not charity. That&#39;s the smart position, held with a longer time horizon than next quarter.&lt;/p&gt;
&lt;h2 id=&#34;redesign-the-role-dont-delete-it&#34;&gt;Redesign the role; don&#39;t delete it&lt;/h2&gt;
&lt;p&gt;The answer isn&#39;t to protect junior work by banning the tools and having bright twenty-three-year-olds do by hand what a model does in seconds. That&#39;s just expensive nostalgia, and they&#39;d hate it. The answer is to redesign what junior means.&lt;/p&gt;
&lt;p&gt;Build the role around judgement instead of rote production. If the model writes the first draft, the junior&#39;s job becomes critiquing it, finding where it&#39;s wrong, improving it — which, done with a good mentor, actually develops judgement &lt;em&gt;faster&lt;/em&gt; than producing slop from scratch ever did. They get to see more cases, across more situations, sooner. Used deliberately, AI could be the best apprenticeship accelerator we&#39;ve had.&lt;/p&gt;
&lt;p&gt;Concretely:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Protect the learning, not the task.&lt;/strong&gt; Ask what a given junior activity was teaching. If a tool removes the activity, you have to deliberately replace the lesson — through review, through harder problems brought forward, through exposure — or you&#39;ve cut the apprenticeship without noticing and kept the headline.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Make seniors teach, and count it as work.&lt;/strong&gt; If AI absorbs the rote, senior time frees up. Aim that freed time at mentoring rather than at simply shipping more. The transfer of judgement from senior to junior is now the scarce, valuable thing, so resource it like it matters and measure people on it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Give juniors real responsibility earlier, with a safety net.&lt;/strong&gt; When the mechanical floor is handled, push them up the stack sooner — judgement calls, design decisions, talking to actual users — with the support to fail safely. Done right, you grow people faster than the old model did.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Hire for trajectory, not just current output.&lt;/strong&gt; A junior who&#39;s slower than the model today but learning fast is an appreciating asset. One who leans on the tool and never builds underlying skill is a liability you won&#39;t spot for years. Hire and develop for the curve.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The organisations that win the next decade won&#39;t be the ones that cut deepest and fastest on entry-level roles. They&#39;ll be the ones who understood that AI changes what junior people should &lt;em&gt;do&lt;/em&gt; without changing the brute fact that you still have to grow your seniors from somewhere. Automate the junior tasks and skip the junior people, and you&#39;re eating your seed corn. It tastes like a saving right up until planting season, when you find you have nothing left to plant.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>Agentic AI in production: the demo lies</title>
      <link>https://partechsystems.com/blog/agentic-ai-production/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/agentic-ai-production/</guid>
      <pubDate>Wed, 18 Mar 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>Agentic AI</category>
      <category>LLMs</category>
      <category>Production</category>
      <category>Reliability</category>
      <description>The demo where an agent books your travel and files your expenses in one fluid run is real. The system you can run for ten thousand customers without a 3am phone call is a different animal entirely.</description>
      <content:encoded><![CDATA[&lt;p&gt;The demo is always beautiful. Someone types a sentence, the agent thinks out loud, calls three tools, books the flight, files the expense, sends a polite Slack message, and the room claps. I&#39;ve sat in that room. I&#39;ve also been the one woken up six weeks later when the same agent quietly double-booked a customer&#39;s order because step four returned an empty string and nobody had decided what an empty string meant.&lt;/p&gt;
&lt;p&gt;That gap — between the demo and the thing you can actually run for paying customers — is where most of the engineering lives, and it&#39;s where the marketing has nothing to say.&lt;/p&gt;
&lt;p&gt;Start with the maths, because the maths is unforgiving. Say each step in your agent&#39;s chain is 95% reliable. Sounds fine. Now chain ten of them. 0.95 to the tenth power is about 0.60. Your beautiful ten-step workflow succeeds three times in five. Chain twenty steps and you&#39;re flipping a coin. Errors don&#39;t add, they compound, and an agent that plans its own steps will happily generate more steps than you expected. The single most common failure I see isn&#39;t a model &#34;hallucinating&#34; — it&#39;s a perfectly reasonable chain that was simply too long to survive its own length.&lt;/p&gt;
&lt;p&gt;So the first instinct people reach for is exactly the wrong one. &#34;Give the agent more autonomy. Let it figure it out.&#34; No. More autonomy means longer chains, more branching, more places for a 95% step to bite you, and far less ability to reason about what the thing will do before it does it. The teams shipping agents that actually hold up in production are doing the opposite. They&#39;re shortening the leash. Fewer steps. Narrower tool sets. Explicit decision points where a human, or at least a deterministic check, gets to say yes before anything irreversible happens.&lt;/p&gt;
&lt;p&gt;That&#39;s the bit that earns its keep: guardrails, approval gates, idempotency, rollback. None of it is glamorous and none of it shows up in a launch video.&lt;/p&gt;
&lt;p&gt;Approval gates first. Anything that spends money, sends an external message, deletes data, or touches a regulated record should pause and ask. Not because the model is stupid, but because the cost of being wrong is asymmetric. A wrong summary costs you a frown. A wrong refund costs you actual pounds and a furious customer. Gate by consequence, not by confidence — confidence scores from these models are not calibrated and you should stop pretending they are.&lt;/p&gt;
&lt;p&gt;Idempotency is the one people forget until it hurts. Agents retry. Frameworks retry. Networks drop and the orchestrator quietly runs your step again. If &#34;create order&#34; runs twice you&#39;ve created two orders. Every tool an agent can call needs an idempotency key or a dedupe check, the same discipline you&#39;d apply to any distributed system, because that&#39;s what this is. An agent is a distributed system with a language model as a very confident, occasionally drunk, scheduler.&lt;/p&gt;
&lt;p&gt;Rollback follows from that. When step seven fails, what undoes steps one through six? If the answer is &#34;nothing, we hope it doesn&#39;t happen,&#34; you don&#39;t have a production system, you have a prototype with good PR. Saga patterns, compensating actions, transactional outboxes — the same toolkit we used for payment systems twenty years ago. The model is new. The failure modes are not.&lt;/p&gt;
&lt;p&gt;Then there&#39;s tool use, which is where the cracks usually show first. The model calls a tool with a malformed argument. The tool returns an error the model has never seen and the model, being a pattern matcher, invents a plausible recovery that&#39;s completely wrong. Or it calls the right tool with stale data because it forgot what it retrieved four steps ago. Tools need strict schemas, validation at the boundary, and error messages written for the model to read — terse, structured, telling it exactly what to do next. Treat your tool layer like a public API being hit by an enthusiastic intern, because functionally it is.&lt;/p&gt;
&lt;p&gt;And none of this is manageable without observability and evals, which I&#39;ll be blunt about: they are not optional, and if your team treats them as a phase-two nicety you will be debugging in production with print statements and prayer. You need to trace every step — inputs, outputs, tool calls, latencies, the lot — because when something goes wrong at step twelve of a forty-step run, &#34;the agent messed up&#34; is not a diagnosis. And you need evals that run a representative set of real tasks on every change, so you know whether last night&#39;s prompt tweak fixed one thing and broke nine others. Models drift, prompts rot, a provider updates something on their end and your success rate quietly drops four points. Without evals you find out from the customer.&lt;/p&gt;
&lt;p&gt;So where do agents genuinely pull their weight today? In bounded, well-instrumented loops where the cost of error is low or fully reversible. Drafting, triage, code that a human reviews, research with citations you can check, classification, routing. Short chains, tight tools, a human in or near the loop. That&#39;s not a consolation prize — that&#39;s a large and genuinely valuable surface, and we ship it.&lt;/p&gt;
&lt;p&gt;Where are they still a liability? Long autonomous chains touching money, irreversible external actions with no gate, anything where you can&#39;t explain after the fact why it did what it did. Putting an unsupervised agent on those is not innovation, it&#39;s a future incident with a launch date.&lt;/p&gt;
&lt;p&gt;My honest read after building these for real customers: the technology is real and the value is real. But the engineering discipline around it lags the marketing by about two years. The vendors are selling you 2028&#39;s autonomy on 2026&#39;s reliability. Build for the reliability you actually have. Short chains, hard gates, idempotent tools, traces on everything, evals on every change. Do that and agents earn their place. Skip it and you&#39;re just automating your incidents.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The fear nobody mentions in the AI rollout meeting</title>
      <link>https://partechsystems.com/blog/managing-ai-anxiety-at-work/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/managing-ai-anxiety-at-work/</guid>
      <pubDate>Tue, 10 Mar 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhavna Ate</dc:creator>
      <category>Change Management</category>
      <category>People</category>
      <category>AI</category>
      <category>Culture</category>
      <description>When AI arrives, everyone in the room is asking the same thing and nobody is saying it out loud. Reassurance that isn&#39;t backed by action makes it worse. Here&#39;s what a humane rollout actually involves, and what it costs you to get wrong.</description>
      <content:encoded><![CDATA[&lt;p&gt;There&#39;s a question in the room at every AI kickoff I&#39;ve ever sat in. It&#39;s never on the agenda. Nobody raises a hand to ask it. But it&#39;s behind every face turned politely towards the slides: &lt;em&gt;is this coming for my job?&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;And while that question is running, nobody is listening to your rollout plan. They&#39;re doing arithmetic about their mortgage. You can present the most elegant adoption roadmap ever built and it will land on people who&#39;ve already mentally checked out to worry about something you didn&#39;t name. The single biggest mistake leaders make with AI is treating it as a tooling rollout when, for the people on the receiving end, it&#39;s an existential one.&lt;/p&gt;
&lt;h2 id=&#34;why-your-reassurance-bounces-off&#34;&gt;Why your reassurance bounces off&lt;/h2&gt;
&lt;p&gt;The instinct is to open with comfort. &#34;AI won&#39;t replace you, it&#39;ll augment you.&#34; &#34;Nobody&#39;s losing their job.&#34; &#34;This frees you up for higher-value work.&#34;&lt;/p&gt;
&lt;p&gt;People have heard those exact words before — about offshoring, about automation, about every reorganisation that ended with someone clearing their desk. The sentences have a track record, and the track record is mixed. So when a leader says them, the room doesn&#39;t hear reassurance. It hears the noise that historically precedes layoffs. Words spent against that history are worse than wasted; they spend down the trust you&#39;ll need later.&lt;/p&gt;
&lt;p&gt;Reassurance is only as good as the action standing behind it. If you say nobody will lose their job and then a team&#39;s headcount shrinks two quarters later through &#34;natural attrition&#34; that you conspicuously don&#39;t backfill, you didn&#39;t tell the truth, and everyone now knows not to believe the next thing you say. People can handle a hard truth told straight. What corrodes a place is a soft lie told to manage them.&lt;/p&gt;
&lt;p&gt;So tell them what&#39;s actually true, including the parts you&#39;d rather not. Be specific about what will change — these tasks are going to be done differently, this kind of work will shrink, here is what we&#39;re asking you to learn. And be specific about what won&#39;t. Vague optimism reads as evasion. Precision, even uncomfortable precision, reads as respect.&lt;/p&gt;
&lt;h2 id=&#34;what-it-costs-when-you-get-this-wrong&#34;&gt;What it costs when you get this wrong&lt;/h2&gt;
&lt;p&gt;Underestimate the human side and the bill comes due in ways that don&#39;t show up on the project plan, because none of them announce themselves.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Attrition you didn&#39;t choose.&lt;/strong&gt; The people with the most options leave first. They&#39;re not waiting to find out how this goes — they&#39;re the ones a competitor will happily take. You&#39;re left with exactly the workforce you didn&#39;t want to over-index on.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Disengagement.&lt;/strong&gt; Far more common than open resistance, and harder to see. People keep showing up and stop trying. They do the minimum, contribute nothing to making the new tools actually work, and let the initiative die of indifference while technically complying with it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sabotage, mostly passive.&lt;/strong&gt; Not dramatic. The colleague who finds the tool &#34;unreliable&#34; and routes around it. The team that documents every flaw and none of the wins. A rollout the organisation is rooting against does not succeed, no matter how good the software is.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Every one of these costs more than slowing down would have. And none of them appear in a status report until they&#39;re entrenched.&lt;/p&gt;
&lt;h2 id=&#34;a-more-humane-way-through-it&#34;&gt;A more humane way through it&lt;/h2&gt;
&lt;p&gt;I&#39;m not arguing for endless feelings sessions. I&#39;m arguing for a rollout that treats people as participants rather than as a population to be managed. In practice:&lt;/p&gt;
&lt;p&gt;Bring people in before the decisions are made, not after. There is a vast difference between change done &lt;em&gt;with&lt;/em&gt; people and change done &lt;em&gt;to&lt;/em&gt; them, and everyone can feel which one they&#39;re in. The people doing the work know where it&#39;ll break and where it&#39;ll genuinely help. Asking them isn&#39;t a courtesy; it produces a better rollout and gives them some footing instead of pure freefall.&lt;/p&gt;
&lt;p&gt;Be concrete about security where you genuinely can be, and straight about uncertainty where you can&#39;t. If roles are genuinely safe, say so and show how. If some roles will change substantially, say that too, and put real resources behind helping people into what&#39;s next — funded time, actual paths, not a course catalogue and good luck. The people who&#39;ll be most affected deserve the earliest, clearest conversation, not the last and vaguest.&lt;/p&gt;
&lt;p&gt;Make managers the front line, and prepare them, because they carry this. The person who absorbs the day-to-day fear isn&#39;t the executive who announced the strategy — it&#39;s the team lead getting the worried questions in the one-to-one. If that manager is anxious and unsupported, they pass the fear straight down, amplified. Look after the managers and a lot of this becomes manageable.&lt;/p&gt;
&lt;p&gt;Then watch the right signals — engagement, the questions people actually ask, where they go silent — and adjust. A rollout you can&#39;t course-correct isn&#39;t a plan, it&#39;s a gamble.&lt;/p&gt;
&lt;p&gt;The fear is rational. People watching a capable machine learn their job are responding sensibly to real uncertainty, and pretending otherwise insults them. You won&#39;t make it disappear with a confident slide, and you shouldn&#39;t try. You earn your way through it — by being straight about what&#39;s happening, by acting in line with what you said, and by treating the people going through the change as the ones most worth listening to about how it should go. Skip that and you don&#39;t avoid the cost. You just pay it later, with interest, in the quality of the people who decide to stay.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>What we actually spend on AI, and where it leaks</title>
      <link>https://partechsystems.com/blog/llm-cost-ops/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/llm-cost-ops/</guid>
      <pubDate>Sat, 14 Feb 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>LLMOps</category>
      <category>FinOps</category>
      <category>Cost</category>
      <category>Operations</category>
      <description>The token price on the pricing page is the part of LLM cost that&#39;s easy to see and least worth obsessing over. The bill that surprises people is everything around the token — retries, evals, human review, context you stuffed in because retrieval was hard. Treat it as unit economics and the surprises stop.</description>
      <content:encoded><![CDATA[&lt;p&gt;A team once showed me a per-request cost they were proud of — a fraction of a cent. I asked what fraction of those requests succeeded first time. They didn&#39;t know. We pulled the numbers: a chunk failed, retried, escalated to a bigger model, then landed on a human reviewer. Their real cost per finished task was many multiples of the number on the slide. They&#39;d been optimising the cheapest thing in the system.&lt;/p&gt;
&lt;p&gt;That&#39;s the whole story of LLM cost. The token price on the pricing page is easy to see and the least worth obsessing over. The bill that surprises people is assembled from everything &lt;em&gt;around&lt;/em&gt; the token: the retries, the evals, the human review, the context you stuffed in because it was simpler than building retrieval. Treat LLM spend like any other unit-economics problem and most of these stop being surprises.&lt;/p&gt;
&lt;h2 id=&#34;cost-per-request-is-the-wrong-metric&#34;&gt;Cost per request is the wrong metric&lt;/h2&gt;
&lt;p&gt;It&#39;s trivial to compute and it tells you almost nothing. It optimises the cheap part and ignores the expensive parts. The metric that matters is &lt;strong&gt;cost per resolved task&lt;/strong&gt; — the fully loaded cost of getting one unit of actual work done, successfully, to a standard you&#39;d accept.&lt;/p&gt;
&lt;p&gt;This isn&#39;t academic. A request costing a fraction of a cent that fails a quarter of the time, triggers a retry, escalates to a larger model, then lands on a human has a resolved-task cost many times the headline number. Optimise per-request and you&#39;ll happily shave the cheap input while the real cost sits in the failure path you aren&#39;t measuring.&lt;/p&gt;
&lt;h2 id=&#34;where-the-spend-actually-leaks&#34;&gt;Where the spend actually leaks&lt;/h2&gt;
&lt;p&gt;Roughly in order of how often I see teams underestimate it:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Retries and fallbacks.&lt;/strong&gt; Every failed call you retry, every escalation to a bigger model when the small one fell short, is real spend that never shows in your per-call estimate. A double-digit retry rate doubles your cost without ever announcing itself.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Context stuffing.&lt;/strong&gt; When retrieval is hard, teams paste the whole document in. It works. It&#39;s also expensive on every single call, forever. The most common leak I find, because it&#39;s invisible until you look at average input token counts.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Prompt bloat.&lt;/strong&gt; System prompts that grow by accretion — every incident adds a clause, nothing&#39;s ever removed. You pay for those tokens on every request, multiplied by volume.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Evaluation.&lt;/strong&gt; Running evals is non-negotiable for quality, but eval traffic is real model traffic. Teams budget production and forget that a serious eval suite can be a meaningful fraction of the production bill.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Human review.&lt;/strong&gt; The most expensive token in the system is a minute of a person&#39;s time. If your design routes a real share of outputs to human review, that — not the API — is your dominant cost, and it scales with volume the worst.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;None of these show up if you only watch the pricing page. All of them show up in cost per resolved task.&lt;/p&gt;
&lt;h2 id=&#34;two-levers-that-actually-move-the-number&#34;&gt;Two levers that actually move the number&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Caching.&lt;/strong&gt; A surprising fraction of production traffic is repetitive or near-repetitive. Caching exact and semantically similar responses, plus prompt caching for stable context, takes load off the model directly. Cheapest win available, and the first thing to instrument.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Model routing.&lt;/strong&gt; Send the cheap, small model first; escalate to the big one only when the task demands it. Most workloads have a long tail of easy cases that don&#39;t need your most capable model. The discipline is measuring the escalation rate and the quality at each tier — so routing cuts cost without silently degrading the resolved-task rate. Done blind, routing just relocates the leak.&lt;/p&gt;
&lt;h2 id=&#34;self-hosting-when-it-wins-and-when-it-doesnt&#34;&gt;Self-hosting: when it wins and when it doesn&#39;t&lt;/h2&gt;
&lt;p&gt;The instinct, once the bill grows, is to self-host an open-weight model and escape API pricing. Sometimes that&#39;s right. Often it isn&#39;t, and the deciding factor is utilisation.&lt;/p&gt;
&lt;p&gt;Self-hosting swaps a variable per-token cost for a largely fixed one: GPUs, the engineering to run them, the on-call to keep them up. That maths works when you have high, steady volume keeping expensive hardware busy. It works badly when your traffic is spiky or modest — you pay for idle GPUs around the clock to serve a few busy hours, and you&#39;ve taken on operational burden the API was absorbing for you.&lt;/p&gt;
&lt;p&gt;So the straight version. Self-hosting beats API pricing when you have sustained high utilisation, a privacy or latency requirement the APIs can&#39;t meet, and the team to operate inference properly. It loses when volume is low or bursty, when you&#39;d run below capacity, or when &#34;we save on tokens&#34; conveniently ignores the salaries of the people now keeping the cluster alive.&lt;/p&gt;
&lt;h2 id=&#34;the-order-of-operations&#34;&gt;The order of operations&lt;/h2&gt;
&lt;p&gt;When the AI bill needs managing, work it in this order:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Instrument cost per resolved task&lt;/strong&gt;, broken down by feature. You can&#39;t manage what you haven&#39;t attributed.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Find the dominant term.&lt;/strong&gt; Tokens, retries, eval, or human review — usually not what people assume.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Attack that term specifically.&lt;/strong&gt; Trim prompts and context if it&#39;s input tokens. Fix quality to cut retries if it&#39;s the failure path. Reduce review load if it&#39;s people.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Add caching and routing&lt;/strong&gt; as standing infrastructure, escalation rate monitored.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Only then&lt;/strong&gt; model self-hosting, and only if utilisation justifies it.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;AI spend is unit economics, not a technology curiosity. Define the unit, attribute the cost, find the dominant term, fix it. The teams that stay solvent here aren&#39;t the ones who negotiated the best token price. They&#39;re the ones who know their cost per resolved task to the cent and can tell you exactly which term is the leak.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>Reskilling isn&#39;t a memo</title>
      <link>https://partechsystems.com/blog/reskilling-teams-for-ai/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/reskilling-teams-for-ai/</guid>
      <pubDate>Fri, 30 Jan 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhavna Ate</dc:creator>
      <category>Reskilling</category>
      <category>People</category>
      <category>AI</category>
      <category>Teams</category>
      <description>Most &#34;reskilling for AI&#34; programmes are a memo, a webinar and a logo on a slide. Real reskilling is slower, costlier and more uncomfortable than that — and the people most likely to be left behind are rarely the ones the programme is designed for.</description>
      <content:encoded><![CDATA[&lt;p&gt;A leadership team announces that the organisation will &#34;reskill the workforce for AI.&#34; A budget line appears. A platform gets bought, the one with ten thousand courses nobody asked for. Everyone is enrolled. A dashboard somewhere lights up green when completion rates climb. Six months on, the work is being done exactly as it was before, except now there&#39;s a slide that says the workforce has been reskilled.&lt;/p&gt;
&lt;p&gt;I&#39;ve watched this play out enough times to recognise the shape of it before it finishes. The tell is that the programme measures whether people watched things, not whether anything changed. And reskilling, the real kind, is not a content-delivery problem. It&#39;s a change-of-behaviour problem, and behaviour is the most expensive thing in any organisation to shift.&lt;/p&gt;
&lt;h2 id=&#34;training-is-not-reskilling&#34;&gt;Training is not reskilling&lt;/h2&gt;
&lt;p&gt;Let me draw the line clearly, because the whole confusion lives here.&lt;/p&gt;
&lt;p&gt;Training is when someone learns what a tool does. Reskilling is when someone&#39;s actual job changes and they&#39;re capable of doing the new version of it. The first is a webinar. The second is months of doing real work differently, with support, while still being expected to deliver.&lt;/p&gt;
&lt;p&gt;A two-hour session on prompting an LLM is training. It is genuinely worth doing and it changes almost nothing on its own, because the moment people return to a workflow, a backlog, and a manager measuring them on the old output, they revert. The pull of how-we&#39;ve-always-done-it is enormous, and a certificate does not counter it. What counters it is changing the work itself — the templates, the definition of done, what gets reviewed, what gets rewarded — so that using the new capability is the path of least resistance rather than an extra thing to remember on top of a full day.&lt;/p&gt;
&lt;p&gt;If your programme didn&#39;t touch the actual workflow, you ran a training programme and called it reskilling. They are not the same line item, and they don&#39;t cost the same.&lt;/p&gt;
&lt;h2 id=&#34;the-realities-nobody-budgets-for&#34;&gt;The realities nobody budgets for&lt;/h2&gt;
&lt;p&gt;Three things get systematically underestimated. I&#39;ll be specific.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Time.&lt;/strong&gt; Reskilling is not free time appearing from nowhere; it&#39;s productive hours spent learning instead of delivering, which means output dips before it recovers. A team genuinely changing how it works will be slower for a stretch. If you haven&#39;t planned for the dip and protected people through it, you&#39;ve designed a programme that punishes the very people taking part — they fall behind on their normal targets while doing the thing you asked.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Managers.&lt;/strong&gt; This is the one that decides everything, and it&#39;s the one most programmes ignore. If a manager doesn&#39;t change what they ask for and how they assess it, nothing downstream of them changes, regardless of how many courses their reports complete. The manager is the gatekeeper of whether the new skill is allowed to be used. Reskill them first, or accept that you&#39;re shouting past the person who controls the day.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The unglamorous middle.&lt;/strong&gt; Not everyone is a curious early adopter who&#39;ll teach themselves over a weekend. Most of an organisation is competent, busy people who&#39;ll engage if it&#39;s made genuinely doable and ignore it if it&#39;s piled on top of everything else. Design for them, not for the enthusiasts who didn&#39;t need you anyway.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&#34;who-gets-left-behind-and-the-deliberate-work-of-not-doing-that&#34;&gt;Who gets left behind, and the deliberate work of not doing that&lt;/h2&gt;
&lt;p&gt;Here&#39;s the uncomfortable part. &#34;Reskilling&#34; is often the polite word organisations use while a cohort of people quietly gets left behind — usually the ones who&#39;ve been doing a now-automatable task for fifteen years, are excellent at it, are in their fifties, and have absorbed an unspoken message that the future doesn&#39;t include them. The programme technically applies to them. The design does not consider them.&lt;/p&gt;
&lt;p&gt;Leaving them behind is rarely a decision anyone makes out loud. It&#39;s what happens by default when you don&#39;t decide otherwise. So you have to decide otherwise, on purpose. That means looking at who is most exposed before you start, building paths that connect their deep domain knowledge to where the work is going rather than treating them as beginners, and being plain with people about what&#39;s changing instead of letting them infer their fate from a vague all-staff email. The domain knowledge in a long-tenured person is often the scarcest asset in the building. Burning it because the retraining was designed for twenty-five-year-olds is an expensive mistake dressed up as progress.&lt;/p&gt;
&lt;h2 id=&#34;did-it-work&#34;&gt;Did it work?&lt;/h2&gt;
&lt;p&gt;Completion rates tell you nothing. Here&#39;s what I&#39;d actually look at, three to six months after the courses end:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Has the work changed? Pull real artefacts — documents, tickets, code, decks — and check whether they look different from before. If they&#39;re identical, the skill isn&#39;t being used, whatever the dashboard says.&lt;/li&gt;
&lt;li&gt;Are people using the capability without being told to? Voluntary, unprompted use is the only honest signal that something stuck.&lt;/li&gt;
&lt;li&gt;Did anyone get left behind, measurably? Look at engagement and attrition in the most-exposed groups specifically. A programme that &#34;succeeded&#34; overall while a cohort silently disengaged did not succeed.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Reskilling done properly is slow, costs real money beyond the platform licence, makes the team temporarily slower, and demands that managers change first. None of that fits neatly on a slide. That&#39;s precisely why so few organisations actually do it, and why the ones that do will be the ones still standing with their people intact on the other side.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>Small models, big deal</title>
      <link>https://partechsystems.com/blog/small-on-device-models/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/small-on-device-models/</guid>
      <pubDate>Thu, 22 Jan 2026 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>Small Models</category>
      <category>On-Device</category>
      <category>Edge</category>
      <category>LLMs</category>
      <description>Everyone&#39;s watching the frontier-model leaderboard. Meanwhile the most useful work I&#39;ve shipped this year runs on an 8B model that fits on a phone and never phones home.</description>
      <content:encoded><![CDATA[&lt;p&gt;Every few weeks another frontier model lands, another trillion parameters, another leaderboard, another round of breathless coverage. Fine. Some of it matters. But the most useful AI work my team has shipped in the last year didn&#39;t touch a frontier model at all. It ran on an 8B model, fine-tuned on one narrow task, quantised down to run on hardware a customer already owned. No data left the building. Latency under a hundred milliseconds. Cost per query rounding to zero.&lt;/p&gt;
&lt;p&gt;Nobody wrote a press release about it. It just worked, and it kept working when the internet didn&#39;t.&lt;/p&gt;
&lt;p&gt;Here&#39;s the bit the size race keeps missing. For an enormous number of real tasks, you don&#39;t need a model that can write a sonnet, debug Rust, and explain Ottoman tax policy. You need a model that does one thing — classify this ticket, extract these fields, redact this document, answer questions about this manual — reliably, cheaply, and fast. A general model the size of a small country is overkill for that, and overkill has costs that don&#39;t show up on the leaderboard.&lt;/p&gt;
&lt;p&gt;The wins from going small are not marginal. They&#39;re structural.&lt;/p&gt;
&lt;p&gt;Privacy first, because it&#39;s the one that closes deals. When the model runs on the device or on a server inside the customer&#39;s own walls, the sensitive data never crosses the network. For anyone in healthcare, finance, defence, or any regulated space, &#34;the data never leaves&#34; isn&#39;t a feature, it&#39;s the entire precondition for the conversation. I&#39;ve watched procurement processes that would&#39;ve taken nine months collapse to weeks the moment we could say the inference happens on-premise. You cannot leak what you never send.&lt;/p&gt;
&lt;p&gt;Latency next. A round trip to a cloud API is a hundred, two hundred milliseconds before the model has thought about anything. On-device, you skip the journey entirely. For anything interactive — a keyboard, a camera, a control loop, a thing a human is waiting on — that difference is the difference between magic and irritating. And it doesn&#39;t fall over when the connection does. Edge devices in a factory, a vehicle, a remote site, a hospital basement with terrible signal — they need to keep working regardless. A model that requires a healthy uplink to function is a model that fails exactly when you need it most.&lt;/p&gt;
&lt;p&gt;Then cost. Frontier inference billed per token is fine until you&#39;re doing it a few million times a day, at which point the bill becomes a strategic problem. A small model on hardware you already paid for has a marginal cost that approaches nothing. At volume that&#39;s not a saving, it&#39;s a different business model.&lt;/p&gt;
&lt;p&gt;Now, the honest part, because going small is not free and anyone who tells you otherwise is selling something. The constraints are real and you have to engineer around them.&lt;/p&gt;
&lt;p&gt;Quantisation is the main lever and the main trap. Dropping from 16-bit to 8-bit weights usually costs you almost nothing measurable and roughly halves your memory. Push to 4-bit and you&#39;ll often still be fine — but &#34;often&#34; is doing real work in that sentence. On some tasks 4-bit quietly degrades in ways that don&#39;t show up until you look closely, and the degradation isn&#39;t uniform across what the model can do. You don&#39;t get to assume. You measure, on your task, with your data.&lt;/p&gt;
&lt;p&gt;Memory is the hard wall. An 8B model in 4-bit is roughly five gigabytes before you&#39;ve loaded a single token of context, and context eats memory too. On a phone, a browser tab, an embedded board, that ceiling decides everything. Half of small-model engineering is just making the thing fit and stay fit while it runs.&lt;/p&gt;
&lt;p&gt;Which is why eval discipline matters more here, not less. With a giant general model you can lean on its slack. A small fine-tuned model has no slack — it&#39;s sharp on the task you trained it for and it falls off a cliff just outside that. So you need a real evaluation set drawn from real inputs, and you need to run it every time you change the model, the quantisation, or the fine-tune. The failure mode I see constantly is a team that fine-tuned once, eyeballed a few outputs, declared victory, and never built the harness that would&#39;ve told them when it broke.&lt;/p&gt;
&lt;p&gt;The payoff, when an 8B model fine-tuned on your task beats a frontier model, is real and more common than people expect. On a narrow, well-defined job with good training data, the small specialist regularly outperforms the giant generalist — and it&#39;s faster, cheaper, and private while doing it. The generalist knows a little about everything. The specialist knows your thing cold.&lt;/p&gt;
&lt;p&gt;So here&#39;s the heuristic I actually use when someone asks what size model to reach for. If the task is narrow, repetitive, and you can describe &#34;right&#34; precisely — start small, fine-tune, and only scale up if the evals demand it. If the task is open-ended, varied, needs broad world knowledge or genuine multi-step reasoning across domains you can&#39;t enumerate — reach for the big general model and pay for it. And if it sits in between, prototype on the big model to prove the task is even solvable, then distil down to the smallest thing that holds the bar. Prove it works, then make it small.&lt;/p&gt;
&lt;p&gt;The frontier will keep climbing and that&#39;s fine. But most of the value in the next few years won&#39;t be at the frontier. It&#39;ll be in thousands of small, sharp, boring models running quietly close to the data. That&#39;s where the engineering is interesting and the economics actually close.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The exit you have never tested</title>
      <link>https://partechsystems.com/blog/model-portability-lock-in/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/model-portability-lock-in/</guid>
      <pubDate>Thu, 11 Dec 2025 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>LLMs</category>
      <category>Architecture</category>
      <category>Open Source</category>
      <category>Risk</category>
      <description>Open weights versus closed weights is usually argued as a values question. It isn&#39;t. It&#39;s a question about switching cost, and almost nobody has measured theirs.</description>
      <content:encoded><![CDATA[&lt;p&gt;The open-weights argument is nearly always conducted as a values debate. Openness good, gatekeepers bad, or alternatively, safety serious, amateurs reckless. I find I have very little to contribute to that, partly because both sides are arguing about the world and I&#39;m usually being paid to argue about one company&#39;s next eighteen months.&lt;/p&gt;
&lt;p&gt;The question I actually care about is duller and more answerable. If the model you&#39;re building on became unavailable, unaffordable, or unacceptable in six months, what would it cost you to move?&lt;/p&gt;
&lt;p&gt;Most teams have never worked this out. Which means they&#39;ve taken on a risk of unknown size, which is the only kind that really hurts.&lt;/p&gt;
&lt;h2 id=&#34;why-the-question-comes-up-now&#34;&gt;Why the question comes up now&lt;/h2&gt;
&lt;p&gt;Three things happen, and they&#39;ve all happened to clients of mine in the past year.&lt;/p&gt;
&lt;p&gt;A hosted model gets deprecated. You&#39;re given a migration window, which sounds generous until you realise your prompts were tuned against the old version&#39;s particular quirks and the replacement is better on benchmarks and worse on your task. That&#39;s not hypothetical; better-on-average routinely means worse-on-yours, and you find out with three weeks left.&lt;/p&gt;
&lt;p&gt;The price changes, or your volume does. At pilot scale nobody cares about per-token cost. At production scale the bill becomes a line item somebody senior has opinions about, and suddenly the question of whether this workload could run somewhere cheaper is not academic.&lt;/p&gt;
&lt;p&gt;Or the requirement changes underneath you. A customer&#39;s procurement asks where the data goes. A regulator asks the same. A deal you want requires processing to stay in a particular jurisdiction, and your entire architecture assumes an API endpoint you don&#39;t control.&lt;/p&gt;
&lt;p&gt;None of these are exotic. They&#39;re the ordinary weather of running software with a supplier in the critical path, and we all knew how to think about this before AI turned up — we just seem to have forgotten while we were excited.&lt;/p&gt;
&lt;h2 id=&#34;what-actually-locks-you-in&#34;&gt;What actually locks you in&lt;/h2&gt;
&lt;p&gt;Here&#39;s the useful part, because the lock-in isn&#39;t where people look for it.&lt;/p&gt;
&lt;p&gt;It isn&#39;t the API. Swapping one chat completion endpoint for another is an afternoon, and there are half a dozen libraries that will do it for you. If your migration plan is &#34;we use an abstraction layer,&#34; you have solved the five percent of the problem that was never going to be the problem.&lt;/p&gt;
&lt;p&gt;The lock-in is in three other places.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Your prompts.&lt;/strong&gt; They are not portable and pretending otherwise wastes a fortnight. A prompt that&#39;s been refined over months against one model encodes hundreds of small accommodations to that model&#39;s tendencies — how it handles a negative instruction, whether it needs the format example, how it drifts when the context gets long. Move it and it degrades in ways that are individually minor and collectively fatal. Expect to redo real work per prompt, and budget for it honestly.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Your fine-tunes.&lt;/strong&gt; If you&#39;ve fine-tuned a hosted model, you have a derivative of an asset you don&#39;t hold. The training data is yours, the artefact isn&#39;t. This is fine if you kept the data, the pipeline, and the ability to run it again. Many teams did not keep the pipeline, because it was run once, by a contractor, in a notebook.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Everything you&#39;ve built around one model&#39;s failure modes.&lt;/strong&gt; The guardrails, the retry logic, the output validators, the little cleanup function that strips the preamble it always adds. That accumulated scar tissue is quietly the most expensive thing to reproduce, and it&#39;s invisible in any architecture diagram.&lt;/p&gt;
&lt;h2 id=&#34;the-thing-that-makes-an-exit-possible&#34;&gt;The thing that makes an exit possible&lt;/h2&gt;
&lt;p&gt;There&#39;s exactly one investment that converts a scary migration into a boring one, and it&#39;s the same artefact I bang on about constantly: a real evaluation set.&lt;/p&gt;
&lt;p&gt;If you have a few hundred graded examples that represent what your system is for, then swapping models is a Tuesday. You point the harness at the candidate, you get numbers, you see precisely where it&#39;s worse, you fix those cases or you decide the trade is acceptable. The decision becomes evidence-based and takes days.&lt;/p&gt;
&lt;p&gt;Without it, the migration is a matter of opinion, conducted under time pressure, by people who are frightened. I&#39;ve watched that go badly enough times to be blunt about it: your evals &lt;em&gt;are&lt;/em&gt; your portability. Everything else is architecture theatre.&lt;/p&gt;
&lt;h2 id=&#34;so-open-or-closed&#34;&gt;So, open or closed?&lt;/h2&gt;
&lt;p&gt;Given all that, my actual advice is unromantic.&lt;/p&gt;
&lt;p&gt;Use the best hosted model for the work that&#39;s hard, novel, and low-volume, where capability matters more than anything and the frontier is genuinely ahead. Pay for it. Don&#39;t self-host to prove a point; running inference well is a real operational discipline and most teams underestimate it by a factor of several.&lt;/p&gt;
&lt;p&gt;Move the high-volume, narrow, well-specified work onto weights you hold, once it&#39;s stable — not for ideology, but because at volume the economics and the latency and the &#34;your data never leaves&#34; conversation all point the same way, and because a model you host doesn&#39;t get deprecated on someone else&#39;s schedule.&lt;/p&gt;
&lt;p&gt;And regardless of which you pick: keep your training data and your pipeline, keep your evals current, and once — just once — actually run the migration. Take a real workload, point it at a different model, and see what breaks. A day spent finding out is cheaper than a quarter spent discovering it during an incident.&lt;/p&gt;
&lt;p&gt;The teams that will handle the next shift calmly aren&#39;t the ones who picked correctly. They&#39;re the ones who know what picking wrongly would cost.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>RAG done right: it&#39;s a search problem, not an AI problem</title>
      <link>https://partechsystems.com/blog/rag-done-right/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/rag-done-right/</guid>
      <pubDate>Mon, 10 Nov 2025 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>RAG</category>
      <category>LLMs</category>
      <category>Search</category>
      <category>Retrieval</category>
      <description>Most RAG projects underperform for the same reason — the team treated retrieval as an AI problem when it was a search problem all along. The model was never your weak link.</description>
      <content:encoded><![CDATA[&lt;p&gt;I&#39;ve now reviewed enough disappointing retrieval-augmented generation projects to spot the diagnosis before anyone opens a laptop. The team is brilliant. The model is the latest and best. The vector database is the fashionable one. And the answers are still wrong, vague, or confidently citing the wrong document. Then someone suggests trying a bigger model, and that&#39;s the moment I know exactly what went wrong.&lt;/p&gt;
&lt;p&gt;They built an AI project. They needed a search project.&lt;/p&gt;
&lt;p&gt;RAG is, mechanically, very simple. You find the relevant bits of your knowledge base, you stuff them into the prompt, the model answers using them. That&#39;s it. And in that pipeline the model is almost never the bottleneck. The bottleneck is the &#34;find the relevant bits&#34; step — the retrieval — which is a search and information-retrieval problem that the field spent thirty years working on before transformers were a twinkle in anyone&#39;s eye. The teams that nail RAG are the ones who remember that. The teams that struggle are the ones who think the embedding model absolves them of doing search properly.&lt;/p&gt;
&lt;p&gt;If the right chunk never makes it into the prompt, the cleverest model on earth will answer from thin air. Garbage retrieval in, confident nonsense out — and now it&#39;s nonsense with a citation, which is worse, because it looks trustworthy.&lt;/p&gt;
&lt;p&gt;So let&#39;s talk about the unglamorous things that actually decide whether your RAG works.&lt;/p&gt;
&lt;p&gt;Chunking, first, because it&#39;s the most underrated decision in the whole stack. How you split your documents determines what can ever be retrieved. Split too small and you shred the context — half a sentence retrieves cleanly and means nothing. Split too big and each chunk is a muddle of five topics, so the embedding represents none of them well and the model drowns in noise. There&#39;s no universal right answer; it depends on your documents. But splitting on real structure — sections, headings, semantic boundaries — beats blindly cutting every 500 tokens almost every time. I&#39;ve seen a RAG system go from useless to genuinely good with no model change at all, purely by fixing how documents were chopped up.&lt;/p&gt;
&lt;p&gt;Metadata, next, and this is where most teams leave the biggest win on the floor. Your chunks have properties — source, date, author, document type, department, version. If you&#39;re not capturing and filtering on those, you&#39;re asking a similarity search to do a librarian&#39;s job. &#34;What&#39;s our current refund policy&#34; should never surface the 2019 version, and no amount of semantic similarity reliably prevents that. A date filter does, instantly. Half the retrieval failures I see are really metadata failures wearing a trench coat.&lt;/p&gt;
&lt;p&gt;Now the one people argue with me about: hybrid beats pure vector far more often than the vector-database marketing admits. Pure semantic search is genuinely good at meaning and genuinely bad at exact matches — product codes, error numbers, proper nouns, acronyms, that one specific term of art your industry uses. Someone searches for error code &#34;E-4471&#34; and a vector search helpfully returns chunks about errors in general. Old-fashioned keyword search — BM25, the technology that ran search engines for decades — nails the exact match every time. Run both, combine the rankings, and you get the meaning and the precision. It&#39;s more engineering than flipping on a vector store, which is precisely why people skip it and then wonder why exact-match queries fail.&lt;/p&gt;
&lt;p&gt;Which brings me to the discipline that separates the projects that improve from the ones that plateau: evaluate retrieval separately from generation. These are two different systems and they fail differently. If you only judge the final answer, you can never tell whether a bad answer came from bad retrieval or bad generation, and so you can&#39;t fix it — you just keep swapping models and changing nothing that matters. Build a set of real questions with the documents that should answer them. Measure whether retrieval surfaces the right chunks — recall, precision, where the right chunk ranks. Get that solid first. Only then worry about how the model phrases things. Fix the search before you touch the prose.&lt;/p&gt;
&lt;p&gt;And under all of it sits the work nobody wants: data hygiene. Duplicate documents, three contradictory versions of the same policy, scanned PDFs that OCR&#39;d into gibberish, tables that flattened into word soup, stale content nobody flagged as dead. RAG over a messy knowledge base produces messy answers with total confidence. The least glamorous and most valuable thing you can do for a RAG project is clean the underlying data — deduplicate, version, fix the extraction, prune the dead weight. It&#39;s tedious and it&#39;s most of the actual job, and the teams that skip it are the teams that ship something embarrassing.&lt;/p&gt;
&lt;p&gt;Last thing, and I mean it: sometimes you don&#39;t need RAG at all, and reaching for it is just complexity cosplay. If the relevant information fits comfortably in the context window, put it in the prompt — context is cheap now and a retrieval pipeline you don&#39;t need is a liability you have to maintain. If the knowledge is stable and bounded, fine-tuning may serve better than retrieving the same things forever. If the answer lives in a database or an API, give the model a tool call and let it fetch the precise answer rather than fuzzy-matching against embedded text. RAG is the right tool for a large, changing corpus of unstructured text you need to ground answers in. It is not a default, and treating it as one is how you end up maintaining a vector database to solve a problem a WHERE clause would&#39;ve handled.&lt;/p&gt;
&lt;p&gt;Get the search right and the AI part mostly takes care of itself. Get the search wrong and no model will save you. That&#39;s the whole lesson, and it predates the hype by about three decades.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>In hiring, the bias you automate is the bias you scale</title>
      <link>https://partechsystems.com/blog/ai-hiring-recruitment/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/ai-hiring-recruitment/</guid>
      <pubDate>Mon, 15 Sep 2025 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhavna Ate</dc:creator>
      <category>Hiring</category>
      <category>People</category>
      <category>AI</category>
      <category>Fairness</category>
      <description>Automate a biased hiring process and you don&#39;t fix the bias. You industrialise it. A look at where AI genuinely helps in recruitment, where it&#39;s dangerous, and who stays accountable when it gets a decision wrong.</description>
      <content:encoded><![CDATA[&lt;p&gt;A recruiter rejecting one CV an hour on a hunch is a small, containable problem. You can coach that person. You can audit their decisions. You can sit them down and ask why every name they shortlisted sounds like their own. A model rejecting ten thousand CVs an hour on a pattern nobody can fully articulate is a different kind of problem entirely, and most of the companies buying these tools haven&#39;t worked out that it&#39;s worse.&lt;/p&gt;
&lt;p&gt;That&#39;s the thing I keep saying in rooms where someone has just demoed a shiny &#34;AI-powered talent platform&#34; and the heads are nodding. Automating a process doesn&#39;t sanitise it. If the process was biased, you&#39;ve now made the bias faster, cheaper, and far harder to see. You&#39;ve taken something a human did clumsily and given it the gloss of mathematical objectivity, which is exactly the disguise discrimination wants.&lt;/p&gt;
&lt;h2 id=&#34;where-the-danger-actually-sits&#34;&gt;Where the danger actually sits&lt;/h2&gt;
&lt;p&gt;I spent two decades inside large technology organisations before I moved into People work, and the engineering habit that stuck with me is this: be specific about which part of the system you&#39;re worried about. &#34;AI in recruitment&#34; is too broad to have an opinion on. So let me split it.&lt;/p&gt;
&lt;p&gt;Ranking and scoring candidates is the dangerous end. A model trained on who your company hired and promoted in the past learns, with great fidelity, who your company hired and promoted in the past. If that history skewed male, or skewed towards one set of universities, or invisibly penalised a two-year gap that usually means childcare, the model absorbs all of it and presents the result as a score out of a hundred. The most cited example is still Amazon&#39;s experimental tool that taught itself to downgrade CVs containing the word &#34;women&#39;s&#34;, but I don&#39;t lean on it because it&#39;s old — I lean on it because it&#39;s typical. The model did exactly what it was built to do. The training data was the discrimination.&lt;/p&gt;
&lt;p&gt;Video &#34;personality&#34; scoring is worse, and I&#39;ll be blunt about it: I think most of it is closer to phrenology than science. Tools that infer conscientiousness or &#34;culture fit&#34; from facial micro-expressions, vocal tone, or word choice are making confident claims about interior states they cannot observe. They penalise accents, neurodivergence, a bad webcam, a candidate interviewing from a noisy flat because they share it with three other people. Illinois and a handful of other places now regulate exactly this for good reason. If you are scoring a human being&#39;s worth from how their face moves on a laptop camera, you have lost the plot.&lt;/p&gt;
&lt;h2 id=&#34;where-it-genuinely-earns-its-place&#34;&gt;Where it genuinely earns its place&lt;/h2&gt;
&lt;p&gt;I&#39;m not against the technology. I&#39;m against pointing it at the decision. Used on the work around the decision, it&#39;s straightforwardly useful.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Scheduling.&lt;/strong&gt; Coordinating five interviewers, two time zones and a candidate&#39;s day job is miserable admin, and there is no fairness question buried in a calendar invite. Automate all of it.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sourcing breadth.&lt;/strong&gt; Search tools that surface candidates you&#39;d never have found — different regions, non-obvious title histories, people who don&#39;t use the keywords your team happens to use — widen the top of the funnel. That&#39;s the opposite of bias, as long as the widening is where it stops and a person decides who to actually approach.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Drafting and tidying.&lt;/strong&gt; First-pass job descriptions, summarising a long CV so a human reads it faster, flagging that a posting is full of jargon that deters the very people you want. Assistive, reversible, low-stakes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The pattern, if you want one line to carry out of here: let the machine widen the funnel and handle the logistics. Don&#39;t let it narrow the field. Breadth is safe. Judgement is not.&lt;/p&gt;
&lt;h2 id=&#34;somebodys-name-has-to-be-on-it&#34;&gt;Somebody&#39;s name has to be on it&lt;/h2&gt;
&lt;p&gt;Here&#39;s the part people skip because it&#39;s inconvenient. For every consequential decision, a named human being has to be accountable — not the vendor, not &#34;the algorithm&#34;, a person you could put in front of a tribunal. The moment a rejected candidate, or a regulator, or your own board asks &#34;why was this person screened out&#34;, the answer cannot be a shrug at a black box. Under the EU AI Act, recruitment systems are high-risk and that accountability is becoming a legal requirement, not a nicety. But I&#39;d want it even if no law demanded it, because a decision nobody will own is a decision nobody can defend.&lt;/p&gt;
&lt;p&gt;So when I deploy any of this, the rules are dull and non-negotiable. The tool recommends; a person decides and signs. We keep records of why, in language a human wrote. We test outcomes across groups before go-live and on a schedule after — not because we expect a clean bill of health, but because the only honest assumption is that bias is present until measured otherwise. We tell candidates an automated step exists and give them a route to a human. And we keep one uncomfortable question pinned to the wall: if this model is wrong about someone, who finds out, and how?&lt;/p&gt;
&lt;p&gt;Most vendors can&#39;t answer that last one. That&#39;s usually all I need to know.&lt;/p&gt;
&lt;p&gt;The mess of it is that good hiring was always hard, slow, and human, and AI is being sold as the thing that finally makes it fast, cheap, and clean. It can make it faster and cheaper. Clean is the lie. If you remember nothing else: a tool that scales your hiring also scales whatever was already wrong with it, and &#34;the system did it&#34; has never once been an acceptable reason to end someone&#39;s chance at a job.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The voice she recorded while she still could</title>
      <link>https://partechsystems.com/blog/ai-accessibility/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/ai-accessibility/</guid>
      <pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Dimple Paratey</dc:creator>
      <category>Accessibility</category>
      <category>AI</category>
      <category>People</category>
      <category>Design</category>
      <description>The most affecting thing I&#39;ve seen AI do this year had nothing to do with productivity. It&#39;s also the area where a good tool and a lazy tool look identical from the outside — and only one of them helps.</description>
      <content:encoded><![CDATA[&lt;p&gt;A friend of a friend — I&#39;ll call her Jo, because I&#39;ve not asked her permission to write this and I&#39;d rather be careful — was diagnosed with motor neurone disease three years ago. One of the things a speech therapist told her early on, at a point when she could still speak more or less normally, was to bank her voice. Sit down and record a few hundred phrases so that when the disease takes her speech, the device she&#39;ll use to talk sounds like her.&lt;/p&gt;
&lt;p&gt;Ten years ago that meant hours of recording in a clinical setting and a synthetic voice that was recognisably hers in the way a photocopy of a photograph is recognisably you. Now it takes a fraction of the recording and the result is close enough that her son, on the phone, hears his mum.&lt;/p&gt;
&lt;p&gt;I&#39;ve thought about that a lot. Not because it&#39;s uplifting — it isn&#39;t, particularly, it&#39;s a terrible illness and a good tool doesn&#39;t change that — but because it&#39;s the clearest example I know of AI doing something that could not be done any other way, for people who were not the target market of anything.&lt;/p&gt;
&lt;h2 id=&#34;the-unglamorous-version-which-is-most-of-it&#34;&gt;The unglamorous version, which is most of it&lt;/h2&gt;
&lt;p&gt;Voice banking is the story that makes people put their glass down at a party. The everyday version matters more, because it affects far more people far more often.&lt;/p&gt;
&lt;p&gt;Live captions that are actually accurate, on any call, without arranging anything in advance. Image descriptions generated on the fly so that a blind person browsing a website isn&#39;t confronted with forty instances of the word &#34;image.&#34; Speech recognition that copes with dysarthria — with a voice that slurs, or shakes, or stops — which mainstream systems were terrible at for years because they were trained on people who sound like newsreaders. Text simplified into plain language for someone with a cognitive disability or a brain injury, on demand, without having to ask a person and feel like a burden.&lt;/p&gt;
&lt;p&gt;Every one of those is a thing that used to require either money, advance planning, or asking another human for help. Removing the &#34;asking another human&#34; step is underrated by people who&#39;ve never had to do it eight times a day.&lt;/p&gt;
&lt;p&gt;There&#39;s an old idea in design about dropped kerbs — cut into pavements for wheelchair users, and then used by everyone with a pram, a suitcase, a delivery trolley. Most of this is a dropped kerb. Live captions are used by disabled people out of necessity and by everyone else on a train with no headphones. The features arrive because of a small group and stay because of a large one, which is usually how the funding gets justified, and I&#39;ve made my peace with that.&lt;/p&gt;
&lt;h2 id=&#34;where-i-get-uneasy&#34;&gt;Where I get uneasy&lt;/h2&gt;
&lt;p&gt;Here&#39;s the part I&#39;d want anyone building this to hear, because it&#39;s the failure mode I keep seeing and it&#39;s a subtle one.&lt;/p&gt;
&lt;p&gt;Automatic accessibility can become an excuse to stop doing accessibility.&lt;/p&gt;
&lt;p&gt;A team turns on AI-generated alt text across their site. The audit score goes up. Somebody screenshots the dashboard. And the alt text now says &#34;a person sitting at a desk&#34; for an image whose entire point was that the chart on the screen shows a 40% drop. Technically described. Functionally useless. Before the tool existed, someone might have written a sentence that mattered. Now nobody will, because the box is ticked and the box is what gets measured.&lt;/p&gt;
&lt;p&gt;Same with captions. Automatic captions are genuinely good now, in good conditions, in a common accent, on ordinary vocabulary. In a lecture on pharmacology, with a lecturer from Newcastle and a room with an echo, they are not good, and the errors are confident and plausible, which is worse than errors that look like errors. A deaf student who can&#39;t tell which words were wrong is in a worse position than one who knows the transcript is unreliable.&lt;/p&gt;
&lt;p&gt;So the test I&#39;d apply is simple, and it&#39;s not a technical test. Does the disabled person get to check it, correct it, and be believed? A system that generates a description and lets the user say &#34;that&#39;s wrong, describe it again, tell me what&#39;s on the screen&#34; is a tool. A system that generates a description, publishes it, and offers no recourse is a compliance product that happens to sit between a person and the thing they were trying to reach.&lt;/p&gt;
&lt;h2 id=&#34;and-the-obvious-thing-which-somehow-needs-saying&#34;&gt;And the obvious thing, which somehow needs saying&lt;/h2&gt;
&lt;p&gt;Ask them.&lt;/p&gt;
&lt;p&gt;The number of accessibility features I&#39;ve seen designed by people who have never watched a screen reader user navigate their product is not small. It&#39;s most of them. An afternoon with four people who actually use this stuff will tell you more than a quarter of internal workshops, and it will spare you shipping something well-meant and infuriating.&lt;/p&gt;
&lt;p&gt;Jo&#39;s voice works, by the way. She says the synthesised version is a bit too polite — that it doesn&#39;t do her swearing properly, which she considers a significant regression. I think that&#39;s the most useful piece of product feedback I&#39;ve heard all year.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The only question I ask in an AI demo</title>
      <link>https://partechsystems.com/blog/how-do-you-know-it-works/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/how-do-you-know-it-works/</guid>
      <pubDate>Wed, 02 Apr 2025 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>Evaluation</category>
      <category>LLMs</category>
      <category>Quality</category>
      <category>Engineering</category>
      <description>There&#39;s one question I ask in every AI demo, and it ends about half of them. Not because the answer is bad — because there isn&#39;t one, and nobody had noticed.</description>
      <content:encoded><![CDATA[&lt;p&gt;I&#39;ve sat through a lot of demos in the past two years and I&#39;ve narrowed my contribution down to one question, usually asked about four minutes in.&lt;/p&gt;
&lt;p&gt;&#34;How do you know it works?&#34;&lt;/p&gt;
&lt;p&gt;Roughly half the time this ends the demo. Not because the answer is bad — because there isn&#39;t one. The team has been iterating for months on a system whose quality they assess by trying it, looking at the output, and going &#34;yeah, that&#39;s better.&#34; Vibes. Very sophisticated vibes, held by clever people with good instincts, but vibes, and vibes don&#39;t survive a version bump.&lt;/p&gt;
&lt;p&gt;Here&#39;s why it matters more than it used to. With conventional software, a regression announces itself. Something throws, a test goes red, a page 500s. With a language model in the loop, a regression is a slightly worse answer, and slightly worse answers look exactly like answers. You can degrade for six weeks and only find out when a customer complains about something that had been quietly broken since the prompt tweak on the 3rd.&lt;/p&gt;
&lt;h2 id=&#34;start-with-the-failures-you-already-have&#34;&gt;Start with the failures you already have&lt;/h2&gt;
&lt;p&gt;The instinct is to build a big representative test set, and that instinct will keep you busy for a month and produce something bland. Don&#39;t start there.&lt;/p&gt;
&lt;p&gt;Start with a folder. Every time the system does something wrong — a bad answer, a hallucinated figure, a refusal that shouldn&#39;t have happened, an escalation that should have been handled — the person who spotted it drops the input, the output, and one line about why it&#39;s wrong into that folder. That&#39;s it. No process, no ticket type, no ceremony, because ceremony is how this dies.&lt;/p&gt;
&lt;p&gt;Within a fortnight you&#39;ll have thirty real failures. Those thirty are worth more than three hundred synthetic examples, because they&#39;re the actual distribution of ways your users break your system, which is never the distribution you&#39;d have imagined. Everything grows from there.&lt;/p&gt;
&lt;p&gt;Somewhere around a hundred and fifty examples you have a genuine evaluation set: mostly real failures, a decent number of cases it handles correctly (so you can catch the fix that breaks something else), and a handful of deliberately nasty edge cases. That set is now the most valuable artefact in the project. More valuable than the prompt, which you will rewrite forty times.&lt;/p&gt;
&lt;h2 id=&#34;get-the-ground-truth-argument-out-of-the-way&#34;&gt;Get the ground truth argument out of the way&lt;/h2&gt;
&lt;p&gt;For each case, someone has to say what the right answer would have been. This is where projects discover that they don&#39;t agree.&lt;/p&gt;
&lt;p&gt;Two of your experts will look at the same output and one will call it acceptable and the other won&#39;t. That disagreement is not an annoyance to be smoothed over; it&#39;s the specification of your product surfacing for the first time. Have the argument. Write down the resolution. If you can&#39;t resolve it, you&#39;ve found a place where your product doesn&#39;t know what it&#39;s for, and no amount of model work will paper over that.&lt;/p&gt;
&lt;p&gt;Where the output is open-ended, don&#39;t try to grade it out of ten. Nobody is consistent at that, including the same person on a different afternoon. Ask narrower binary questions: did it use only the supplied sources? Did it get the number right? Did it refuse when it should have? Four crisp yes/no checks tell you far more than one holistic score, and they tell you &lt;em&gt;what&lt;/em&gt; broke rather than just that something did.&lt;/p&gt;
&lt;h2 id=&#34;on-using-a-model-to-grade-a-model&#34;&gt;On using a model to grade a model&lt;/h2&gt;
&lt;p&gt;You&#39;ll want to automate this, and you can, within limits. Using a strong model as a judge works decently for constrained questions of the kind above. It works poorly for &#34;is this good,&#34; which is unsurprising.&lt;/p&gt;
&lt;p&gt;But calibrate it before you trust it. Take fifty cases you&#39;ve graded by hand, run the judge over them, and see how often it agrees with you. If it&#39;s agreeing 90% of the time on your judgements, you have a useful instrument. If it&#39;s at 70%, you have a random number generator with a nice interface, and any improvement it reports is inside its own noise. I&#39;ve seen teams celebrate a four-point gain from a judge whose agreement rate meant four points was indistinguishable from nothing.&lt;/p&gt;
&lt;p&gt;Also, and this catches people: don&#39;t have the same model family grade its own output on style. It likes its own writing. So would you.&lt;/p&gt;
&lt;h2 id=&#34;make-it-a-gate-not-a-report&#34;&gt;Make it a gate, not a report&lt;/h2&gt;
&lt;p&gt;The last step is the one that changes behaviour. The eval has to run automatically on every change to the prompt, the retrieval, the model version, the temperature, the chunking — every knob — and it has to block the change if the numbers drop. Not email a dashboard. Block.&lt;/p&gt;
&lt;p&gt;The moment that&#39;s in place, a whole class of argument disappears from the team. &#34;Does this prompt change help?&#34; stops being a matter of opinion and seniority and becomes a number that took eight minutes to produce. That shift is worth more than the eval itself, honestly. It&#39;s the difference between a team that improves and a team that oscillates.&lt;/p&gt;
&lt;p&gt;None of this is exotic. It&#39;s the testing discipline that ordinary software worked out decades ago, applied to a component that fails softly instead of loudly. The reason it gets skipped is that it isn&#39;t fun, it produces no demo, and it makes visible how often the system is wrong — which is uncomfortable in month two and priceless in month nine.&lt;/p&gt;
&lt;p&gt;If you can&#39;t answer the question, you don&#39;t have a product yet. You have a very persuasive prototype, and those are much easier to build than they used to be.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The meeting has a memory now</title>
      <link>https://partechsystems.com/blog/ai-notetakers-meetings/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/ai-notetakers-meetings/</guid>
      <pubDate>Thu, 13 Feb 2025 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhavna Ate</dc:creator>
      <category>Meetings</category>
      <category>People</category>
      <category>AI</category>
      <category>Culture</category>
      <description>Meetings got a permanent memory almost overnight and nobody ran a consultation about it. The productivity case is real. So is the quieter thing it does to how honestly people speak.</description>
      <content:encoded><![CDATA[&lt;p&gt;Somewhere in the last eighteen months, without a decision being taken anywhere, the meeting acquired a permanent memory.&lt;/p&gt;
&lt;p&gt;Nobody ran a consultation. There was no policy. Someone connected a notetaker to their calendar because it was useful, a colleague saw the summary and asked what it was, and within a few months it was in every recurring call in the organisation, sitting quietly in the participant list. I&#39;ve watched this happen at four or five clients now and the sequence is almost identical each time.&lt;/p&gt;
&lt;p&gt;I want to be fair to it, because the benefits are not imaginary. And then I want to talk about the part that people feel before they can articulate it.&lt;/p&gt;
&lt;h2 id=&#34;the-genuinely-good-part&#34;&gt;The genuinely good part&lt;/h2&gt;
&lt;p&gt;For a lot of people this technology is the best thing to happen to their working week in years.&lt;/p&gt;
&lt;p&gt;If you have a hearing impairment, a live transcript is not a convenience, it&#39;s access. If English is your third language and the call is six British people talking over each other, being able to read it back at your own pace changes what you can contribute. If you have ADHD and the effort of taking notes competes directly with the effort of listening, being relieved of one of those is enormous. If you&#39;re on parental leave or in a different timezone or simply double-booked, catching up in four minutes rather than being permanently half-informed is the difference between being in the loop and drifting out of it.&lt;/p&gt;
&lt;p&gt;I don&#39;t think any of that gets said enough, because the discourse about meeting notetakers is mostly written by people for whom meetings were already working fine.&lt;/p&gt;
&lt;p&gt;There&#39;s also the ordinary version: actions get captured, the &#34;what did we agree?&#34; argument three weeks later gets settled in ten seconds, and nobody has to nominate the most junior person in the room to type while everyone else thinks.&lt;/p&gt;
&lt;h2 id=&#34;what-changes-in-the-room&#34;&gt;What changes in the room&lt;/h2&gt;
&lt;p&gt;Here&#39;s the part I&#39;d want any leadership team to sit with.&lt;/p&gt;
&lt;p&gt;When people know the room has a permanent, searchable, exportable memory, they speak differently. Not dramatically. Nobody announces that they&#39;ve become more careful. But the half-formed idea doesn&#39;t get floated. The honest &#34;I don&#39;t actually understand this&#34; gets swallowed. The bit where someone says &#34;look, between us, the timeline is fantasy&#34; — that stops, and it stops first among the people with the least security, which is precisely the people you most need to hear from.&lt;/p&gt;
&lt;p&gt;I asked a group at one client, anonymously, whether the notetaker had changed how they talked in meetings. Just over half said yes. Of those, almost all described the change as being more guarded. Two people said it made them more likely to speak up because they knew they&#39;d be credited properly, which I thought was an interesting and genuine counterweight, but the balance was clear.&lt;/p&gt;
&lt;p&gt;Psychological safety is not a soft concept. It is the mechanism by which problems reach you early enough to be cheap. If a tool trades some of that for tidier action items, that&#39;s a trade you should make consciously, and most organisations have made it accidentally.&lt;/p&gt;
&lt;h2 id=&#34;the-summary-is-not-the-meeting&#34;&gt;The summary is not the meeting&lt;/h2&gt;
&lt;p&gt;Second thing, more mundane and probably more immediately damaging: these summaries are confidently wrong in specific, recurring ways, and people treat them as the record.&lt;/p&gt;
&lt;p&gt;They flatten disagreement. A discussion where three people had reservations and one senior person pushed through becomes &#34;the team agreed to proceed.&#34; They mis-attribute. They struggle with the person who said the crucial thing quietly at minute fifty-two after everyone had mentally moved on. They cannot hear irony, and British professional life runs substantially on irony.&lt;/p&gt;
&lt;p&gt;None of that would matter if the summary were treated as a convenience. It matters because within about a month the summary becomes the thing people cite, and the meeting itself becomes something nobody remembers independently. You&#39;ve created an authoritative record with a systematic bias toward consensus and seniority, and no named human is accountable for its accuracy.&lt;/p&gt;
&lt;h2 id=&#34;what-id-actually-put-in-place&#34;&gt;What I&#39;d actually put in place&lt;/h2&gt;
&lt;p&gt;I&#39;d resist the instinct to ban it. Bans on useful things produce shadow use, and you lose the access benefits along with the risks.&lt;/p&gt;
&lt;p&gt;Four things, and they fit on one page.&lt;/p&gt;
&lt;p&gt;Say clearly which meetings are recorded by default and which are not. One-to-ones, anything about performance, anything about someone&#39;s health or circumstances, and any conversation whose purpose is to surface bad news — these should be off unless both parties actively want it on. Make &#34;off&#34; socially easy: if declining requires a person to make a small speech, you have not given them a choice.&lt;/p&gt;
&lt;p&gt;Name an owner for every summary. A person, not a tool. They read it before it circulates and they fix what&#39;s wrong. This takes two minutes and it&#39;s the difference between a record and a rumour.&lt;/p&gt;
&lt;p&gt;Be explicit about retention and access. Who can search these? For how long? Can they be pulled into a grievance, a performance case, a redundancy consultation? People will assume the worst answer if you don&#39;t give them the real one, and quite often the worst answer is the true one, in which case they deserve to know.&lt;/p&gt;
&lt;p&gt;And once a quarter, ask people whether it&#39;s changed how they speak. Anonymously. Then actually look at the answer.&lt;/p&gt;
&lt;p&gt;The technology here is not the hard bit. The hard bit is that we installed a memory into the one place where organisations do their honest thinking, and we did it without deciding to.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>The data was always the project</title>
      <link>https://partechsystems.com/blog/data-quality-is-the-project/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/data-quality-is-the-project/</guid>
      <pubDate>Tue, 19 Nov 2024 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>Data</category>
      <category>AI</category>
      <category>Delivery</category>
      <category>Engineering</category>
      <description>Every stalled AI project I&#39;ve been called into has had the same root cause, and it was never the model. It was that nobody wanted to fund the boring six weeks at the start.</description>
      <content:encoded><![CDATA[&lt;p&gt;I get called in when things have stalled. It&#39;s a decent chunk of what I do, and after enough of these I&#39;ve stopped bothering with a diagnostic phase, because the answer is nearly always the same and I can usually confirm it in an afternoon.&lt;/p&gt;
&lt;p&gt;The model is fine. The team is good. Nobody funded the six weeks of unglamorous data work at the start, because six weeks of data work has never once made a board paper look exciting, and now they&#39;re nine months in, spending money at a rate that requires a story, and the story has to be about the model because that&#39;s the part anyone can be persuaded to care about.&lt;/p&gt;
&lt;p&gt;Let me be specific about what actually goes wrong, because &#34;data quality&#34; as a phrase is so worn out that it no longer transmits anything.&lt;/p&gt;
&lt;h2 id=&#34;we-have-loads-of-data&#34;&gt;&#34;We have loads of data&#34;&lt;/h2&gt;
&lt;p&gt;I hear this in the first meeting almost every time, and I&#39;ve learned to treat it as a warning rather than a reassurance.&lt;/p&gt;
&lt;p&gt;Having loads of data means you have loads of rows. It doesn&#39;t tell you whether those rows describe the thing you want to predict, whether they were collected consistently, whether the process that generated them is the same process running today, or whether half the useful signal was thrown away by a form redesign in 2021 that nobody thought to mention.&lt;/p&gt;
&lt;p&gt;A manufacturer once told me they had eight years of sensor readings. True. What they didn&#39;t have was eight years of &lt;em&gt;comparable&lt;/em&gt; sensor readings, because the line had been re-instrumented twice and the units changed the second time. Roughly a third of the archive was measuring something subtly different from the rest, and no field in the schema recorded that. It took two engineers eleven days to work this out, mostly by finding a retired maintenance supervisor who remembered.&lt;/p&gt;
&lt;p&gt;That&#39;s the shape of it. The problem is almost never volume. It&#39;s that the meaning of the data changed over time and nothing wrote it down.&lt;/p&gt;
&lt;h2 id=&#34;nobody-agrees-what-the-label-means&#34;&gt;Nobody agrees what the label means&lt;/h2&gt;
&lt;p&gt;This is the one that quietly kills classification projects.&lt;/p&gt;
&lt;p&gt;You want a model to flag &#34;high-risk&#34; claims, or &#34;urgent&#34; tickets, or &#34;defective&#34; parts. So you get historical examples labelled by the people who do the job. Fine. Then you take a hundred of those examples, hand them to three experienced people independently, and ask them to label them again from scratch.&lt;/p&gt;
&lt;p&gt;Do this. Please do this before you build anything. In my experience, on genuinely hard categories, three experts agree on maybe 60 to 75% of cases, and the disagreements are not random — they&#39;re structural. One person thinks &#34;urgent&#34; means the customer is angry, another thinks it means there&#39;s a contractual clock running. Both have been labelling tickets for years, and the training set is a blend of two incompatible definitions.&lt;/p&gt;
&lt;p&gt;No model can resolve that. It&#39;ll learn the blend, produce mush, and you&#39;ll blame the architecture. The fix is a couple of days in a room with the people who know, arguing until the definition is written down in a way that survives contact with an edge case. This is the least technical work in the whole project and it has the highest leverage of anything you&#39;ll do.&lt;/p&gt;
&lt;h2 id=&#34;the-plumbing-nobody-owns&#34;&gt;The plumbing nobody owns&lt;/h2&gt;
&lt;p&gt;Then there&#39;s access, which stalls more pilots than any modelling problem I&#39;ve seen.&lt;/p&gt;
&lt;p&gt;The data lives in four systems. One of them is a supplier&#39;s, and the contract doesn&#39;t clearly permit this use. One of them is exportable only through a report designed for humans, which means a CSV where the header row is a merged cell and the totals are inline. One of them has a nightly job that everybody assumes is fine because it hasn&#39;t alerted, and which has been silently dropping records with non-ASCII names since the last upgrade. And the person who understands the fourth one left in March.&lt;/p&gt;
&lt;p&gt;None of this is interesting and all of it is real. When a project timeline says &#34;week 1–2: data acquisition,&#34; what it actually means is &#34;weeks 1–9, contingent on a legal review nobody has requested yet.&#34;&lt;/p&gt;
&lt;h2 id=&#34;what-id-do-instead&#34;&gt;What I&#39;d do instead&lt;/h2&gt;
&lt;p&gt;Spend the first six weeks on the data and say so out loud, in the plan, with the money attached, before anyone gets excited. Frame it as what it is: the part that determines whether the rest works.&lt;/p&gt;
&lt;p&gt;In that time, do four things. Pull a real sample and look at it with your own eyes — actual rows, on a screen, not summary statistics, because summary statistics hide exactly the corruption you&#39;re hunting for. Reconstruct the provenance of each field: who enters it, when, under what pressure, and what changed. Run the labelling agreement exercise. And build the smallest end-to-end pipeline you can, moving real data from source to output with nothing clever in the middle, so the access problems surface in week two instead of month five.&lt;/p&gt;
&lt;p&gt;If the data doesn&#39;t survive that, you&#39;ve learned it for the price of six weeks rather than nine months, and you can go and fix the collection process, which was probably the real project anyway.&lt;/p&gt;
&lt;p&gt;There&#39;s a version of this article that ends on something inspiring. I don&#39;t have one. The work is dull, it&#39;s most of the job, and the teams who do it ship while the teams who skip it hold a lot of steering group meetings about model selection.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>Eleven staff, forty thousand case notes</title>
      <link>https://partechsystems.com/blog/ai-charities-nonprofits/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/ai-charities-nonprofits/</guid>
      <pubDate>Tue, 24 Sep 2024 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Dimple Paratey</dc:creator>
      <category>Charities</category>
      <category>AI</category>
      <category>People</category>
      <category>Community</category>
      <description>A woman running a homelessness charity with eleven staff asked me a question I couldn&#39;t answer with any of my usual material. Small charities are where AI either does something genuinely good or wastes money that was donated for something else.</description>
      <content:encoded><![CDATA[&lt;p&gt;I ended up sitting next to a woman called Ruth at a fundraising thing last spring, the sort of evening where everyone has a lanyard and the wine is warm. She runs a homelessness charity in the Midlands. Eleven staff. Somewhere north of forty thousand case notes going back nineteen years, most of them typed into a system that was modern when Blair was in office.&lt;/p&gt;
&lt;p&gt;She asked me, quite directly, whether &#34;all this AI&#34; was something she should be paying attention to or whether it was for organisations with more money than hers. And I found I couldn&#39;t answer with any of my usual material, because everything I normally say assumes a company that can afford to be wrong once.&lt;/p&gt;
&lt;p&gt;Charities can&#39;t. Every pound they spend on a tool that doesn&#39;t work is a pound somebody donated believing it would go somewhere else. That changes the maths, and it changes what a responsible recommendation looks like.&lt;/p&gt;
&lt;h2 id=&#34;where-it-actually-helps&#34;&gt;Where it actually helps&lt;/h2&gt;
&lt;p&gt;I went away and did some homework, and then I went back to Ruth. What I found is that the useful stuff in the third sector is almost embarrassingly mundane, which is exactly why it doesn&#39;t get written about.&lt;/p&gt;
&lt;p&gt;Grant applications. A small charity&#39;s fundraising is often one person writing similar-but-not-identical applications over and over, each with its own word limits and its own idea of what an &#34;outcome&#34; is. The source material — what the charity does, who it helps, what happened last year — barely changes. That&#39;s a genuinely good fit: drafting from material you already have, with a human who knows the funder doing the final pass. Ruth&#39;s fundraiser reckoned it gave her back most of a day a week, and a day a week is a meaningful fraction of an eleven-person organisation.&lt;/p&gt;
&lt;p&gt;Reporting. Funders want impact reports. Impact reports are usually a person pulling numbers out of a case management system and writing the same twelve paragraphs in a slightly different order. Same logic applies.&lt;/p&gt;
&lt;p&gt;And then the case notes. Nineteen years of them. Ruth&#39;s team had a real and rather sad problem: somebody would come through the door for the third time in six years, and unless the worker on duty happened to have been there for all three, that history was effectively invisible. Not lost — sitting right there in the database — just not findable by anyone in the ninety seconds you have before someone stops trusting you. Being able to ask a question of nineteen years of records in plain English is, for that work, not a productivity gain. It&#39;s the difference between treating someone as a stranger and treating them as a person you&#39;ve met before.&lt;/p&gt;
&lt;h2 id=&#34;where-id-be-careful-and-i-mean-properly-careful&#34;&gt;Where I&#39;d be careful, and I mean properly careful&lt;/h2&gt;
&lt;p&gt;That last one is also the one that worries me most, and I told her so.&lt;/p&gt;
&lt;p&gt;Case notes about vulnerable people are about as sensitive as data gets. Names, addresses, mental health, immigration status, children, abuse. If any of that ends up in a system where it&#39;s being used to train someone&#39;s model, or sitting on a server in a jurisdiction nobody checked, the charity has done real harm to people who had very little bargaining power to begin with. That&#39;s not a compliance risk, it&#39;s a betrayal, and it lands on people who came in asking for help.&lt;/p&gt;
&lt;p&gt;So my honest advice to Ruth was: do the grant writing and the reporting first, because the material is public-facing and the downside of getting it wrong is an awkward paragraph. Do the case notes second, slowly, with someone who understands data protection actually in the room, and probably with something that runs on infrastructure you control rather than a free tier that pays for itself with your data.&lt;/p&gt;
&lt;p&gt;The other thing I&#39;d flag, gently, is that the sector is being circled at the moment. Charities are being sold AI transformation by consultancies who&#39;ve noticed that &#34;digital transformation&#34; money exists in the grant landscape. Some of that is fine. Some of it is a discovery phase and a report, funded by a grant, that produces nothing anyone uses. If a proposal doesn&#39;t end with a specific person doing a specific task differently on a specific Tuesday, it isn&#39;t a project, it&#39;s a document.&lt;/p&gt;
&lt;h2 id=&#34;the-bit-that-stayed-with-me&#34;&gt;The bit that stayed with me&lt;/h2&gt;
&lt;p&gt;What struck me most about the conversation wasn&#39;t the technology. It was that Ruth&#39;s instinct was already right and she just didn&#39;t trust it, because everything written about AI is pitched at organisations fifty times her size and written in a register designed to make her feel behind.&lt;/p&gt;
&lt;p&gt;She isn&#39;t behind. She has a clear-eyed view of what her staff spend their time on, which is more than most FTSE boards can say. That&#39;s the whole prerequisite. The tooling question is downstream of it and much less interesting.&lt;/p&gt;
&lt;p&gt;If you&#39;re running something small and worthy and you&#39;re being told you need an AI strategy — you probably don&#39;t. You might need one afternoon, a list of the three things that eat your team&#39;s week, and someone honest enough to tell you when the answer is no.&lt;/p&gt;
&lt;p&gt;We do that conversation for free for charities, as it happens. Ruth&#39;s still going. The case notes project is on hold until they can afford to do it properly, which I think was the right call.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>Build, buy, or wait: the third door</title>
      <link>https://partechsystems.com/blog/build-buy-or-wait/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/build-buy-or-wait/</guid>
      <pubDate>Tue, 06 Aug 2024 00:00:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Bhaskar Paratey</dc:creator>
      <category>Strategy</category>
      <category>Procurement</category>
      <category>AI</category>
      <description>Build-versus-buy is a two-way choice because two-way choices fit on a slide. There&#39;s a third door, and it&#39;s the one that saves the most money and gets picked the least.</description>
      <content:encoded><![CDATA[&lt;p&gt;Every few weeks somebody asks me whether they should build their AI capability in-house or buy it from a vendor. I always answer, because it&#39;s a fair question, but I&#39;ve come to think it&#39;s the wrong shape. Build-versus-buy gets framed as a two-way choice mainly because two-way choices fit on a slide and go to a vote cleanly. There are three doors. The third one is &#34;wait,&#34; and in my experience it is simultaneously the cheapest option available and the one almost nobody picks, because nobody has ever been promoted for recommending it.&lt;/p&gt;
&lt;p&gt;Let me take them in order, since the reasoning is different in each case.&lt;/p&gt;
&lt;h2 id=&#34;buying-and-actually-meaning-it&#34;&gt;Buying, and actually meaning it&lt;/h2&gt;
&lt;p&gt;Buy when the problem you have is a problem thousands of other companies also have. Transcription. Document extraction. Support ticket routing. Meeting summaries. There is no strategic advantage waiting for you at the end of building your own transcription service, and a vendor who does nothing else will be better at it than you within a quarter.&lt;/p&gt;
&lt;p&gt;The failure mode here isn&#39;t buying. It&#39;s buying and then refusing to accept what you bought. A team purchases a platform, discovers it does the job 80% the way they&#39;d like, and then spends fourteen months and a small fortune customising the remaining 20% until they&#39;ve built a bespoke system they can&#39;t upgrade, can&#39;t support, and don&#39;t own the source code to. You&#39;ve now paid the full cost of building plus a licence fee. If a product doesn&#39;t fit your process, the honest options are to change your process or to walk away, and changing the process is usually right and always unpopular.&lt;/p&gt;
&lt;p&gt;Three questions I&#39;d want answered before signing anything. Where does our data go, and can we get it back in a usable form if we leave? Which model is under the hood, and what happens the day you swap it for a cheaper one without telling us? And what does the price look like at ten times our current volume — not because we&#39;ll get there, but because I want to see whether the pricing was designed by someone who thought about it.&lt;/p&gt;
&lt;h2 id=&#34;building-with-your-eyes-open&#34;&gt;Building, with your eyes open&lt;/h2&gt;
&lt;p&gt;Building is right when the thing is genuinely yours — when it runs on data nobody else has, encodes judgement specific to your business, or is the actual product you sell. Fair enough. But be clear about what you&#39;re signing up for, because the initial build is the smallest line item in the whole undertaking.&lt;/p&gt;
&lt;p&gt;What you&#39;re really committing to is a second year. The retraining pipeline when the data drifts. The evaluation harness that tells you whether last week&#39;s change made things better or quietly worse. The monitoring, the on-call rota, the runbook. The awkward Tuesday when the one engineer who understood the whole thing hands in their notice and you discover the design lived entirely in their head. I&#39;ve seen more AI systems die of maintenance neglect than of bad modelling. The first version is a sprint; everything after it is a standing cost, and it doesn&#39;t stop.&lt;/p&gt;
&lt;p&gt;So the test isn&#39;t &#34;can we build this.&#34; Almost any competent team can build a first version now; the tooling is extraordinary. The test is whether you&#39;re prepared to fund it in year three, when it&#39;s boring and works and nobody&#39;s excited about it any more.&lt;/p&gt;
&lt;h2 id=&#34;the-third-door&#34;&gt;The third door&lt;/h2&gt;
&lt;p&gt;Wait when the problem is real but the ground is moving faster than your procurement process. This is more common than people admit. If a category of capability is getting roughly twice as good and half as expensive every year or so, a three-year contract signed today locks you to today&#39;s economics for the entire period in which those economics are going to change most. Sometimes the correct move is to sit still for six months and buy the thing that doesn&#39;t exist yet, cheaper.&lt;/p&gt;
&lt;p&gt;The critical part: waiting is not the same as doing nothing, and if you treat it as an excuse to do nothing you&#39;ve simply lost six months. Waiting well means spending that time on the parts that don&#39;t expire. Write down how the process actually works today, in real detail, including the exceptions people handle by instinct and have never documented. Measure the baseline — how long the task takes now, how often it&#39;s wrong now — because without that number you will never be able to prove any tool helped. Clean up the data. Fix the permissions mess. Sort out who owns which system.&lt;/p&gt;
&lt;p&gt;None of that gets thrown away when you eventually build or buy. All of it is required either way. And most of it is what stalls projects six weeks in, when the pilot grinds to a halt because nobody can get access to the CRM export.&lt;/p&gt;
&lt;h2 id=&#34;the-test-compressed&#34;&gt;The test, compressed&lt;/h2&gt;
&lt;p&gt;I ask three things. Is this differentiating — would a competitor with the identical tool beat us anyway on something else? Is the hard part the data, and is that data ours and messy? And how fast is this category moving right now?&lt;/p&gt;
&lt;p&gt;Generic problem, mature category: buy it, and don&#39;t customise it into a coffin. Differentiating problem, your data, stable enough tooling: build it, and budget for year three. Real problem, category in flux: wait deliberately, and spend the wait on the plumbing.&lt;/p&gt;
&lt;p&gt;The companies I&#39;ve watched get the most out of AI over the last few years weren&#39;t the fastest movers. They were the ones who knew which of their problems was which.&lt;/p&gt;]]></content:encoded>
    </item>
    <item>
      <title>Smart cities are easy to sell. Who are they actually for?</title>
      <link>https://partechsystems.com/blog/ai-urban-planning/</link>
      <guid isPermaLink="true">https://partechsystems.com/blog/ai-urban-planning/</guid>
      <pubDate>Mon, 01 Jul 2024 10:15:00 +0000</pubDate>
      <dc:creator xmlns:dc="http://purl.org/dc/elements/1.1/">Dimple Paratey</dc:creator>
      <category>Cities</category>
      <category>AI</category>
      <category>Surveillance</category>
      <category>People</category>
      <description>There is a man I see most mornings who navigates our high street in a wheelchair, and the city was plainly not built with him in mind. &#34;Smart city&#34; pitches rarely start with him. They should — because the question that matters isn&#39;t how clever the sensors are, it&#39;s who the city is actually for.</description>
      <content:encoded><![CDATA[&lt;p&gt;There&#39;s a man I pass most mornings on the school run. He gets around our high street in a wheelchair, and watching him do it is a daily reminder that the city was designed by people who never had to. The dropped kerb that isn&#39;t quite dropped. The &#34;accessible&#34; route that adds twenty minutes. The pavement café that&#39;s colonised the only flat bit of path.&lt;/p&gt;
&lt;p&gt;I think about him whenever I&#39;m shown a &#34;smart city&#34; pitch, because he is almost never in it. The slides are full of glowing dashboards and sensors and traffic flowing like water, and somewhere in there, unspoken, is a fantasy of a citizen who is young, able-bodied, owns a smartphone, and has nothing to hide. He is not that citizen. Nor, on a bad day, am I.&lt;/p&gt;
&lt;p&gt;&#34;Smart city&#34; isn&#39;t really a technology. It&#39;s a shopping basket — sensors, dashboards, traffic systems, and, far too often, facial recognition — sold to councils as one tidy bundle. Some of what&#39;s in the basket would genuinely make life better for the man on the high street. Some of it would make him a data point whether he likes it or not. The whole job is telling those apart before someone signs a fifteen-year contract.&lt;/p&gt;
&lt;h2 id=&#34;the-stuff-that-genuinely-helps-real-people&#34;&gt;The stuff that genuinely helps real people&lt;/h2&gt;
&lt;p&gt;Let me start with what I&#39;d happily see my own council buy, because it&#39;s worth saying that some of this is properly good.&lt;/p&gt;
&lt;p&gt;Traffic signals are the obvious one. Most of them were timed decades ago using rules of thumb, and they&#39;ve been faithfully causing jams ever since. Adjusting them to what&#39;s actually on the road eases congestion and clears the air a bit — and, crucially, it works perfectly well without knowing who&#39;s sitting in any of the cars. That distinction matters more than anything else in this piece, so hold onto it.&lt;/p&gt;
&lt;p&gt;Then the unglamorous plumbing: predicting which water main or stretch of road is about to fail and fixing it before it floods someone&#39;s street. Tuning the heating in schools and hospitals so they stop burning money. Working out where the ambulances should wait so they reach people faster. None of this makes a flashy demo. All of it changes someone&#39;s actual day.&lt;/p&gt;
&lt;p&gt;And the one closest to my heart — routing that avoids the steps, the broken pavements, the dark underpass. Nobody puts accessibility tech on the conference poster. But for the man on my high street, that&#39;s not a feature. That&#39;s whether he can get to the shops alone.&lt;/p&gt;
&lt;p&gt;Notice what every one of those has in common. A clear job, a saving you can point at, and no need to know who anyone is.&lt;/p&gt;
&lt;h2 id=&#34;the-stuff-id-fight-before-its-installed&#34;&gt;The stuff I&#39;d fight before it&#39;s installed&lt;/h2&gt;
&lt;p&gt;Now the basket gets darker, and the danger is that the cost never shows up on the day you buy it. It shows up years later, on someone else.&lt;/p&gt;
&lt;p&gt;The camera that smooths the traffic can run facial recognition with a software update. The sensor listening for noise can pick out an individual. Once that hardware is bolted to the lampposts, deciding &lt;em&gt;not&lt;/em&gt; to use it that way stops being a setting and becomes a political brawl — one that residents usually lose. You have to draw that line before the kit goes up, because afterwards the kit decides.&lt;/p&gt;
&lt;p&gt;Vendor lock-in is the boring trap that bites hardest. Sign with the wrong supplier and you can lose the right to switch, to open your own data, or to change your mind. That&#39;s not a software preference. That&#39;s handing a private company a chunk of your city&#39;s nervous system.&lt;/p&gt;
&lt;p&gt;And then the systems that decide where to send police patrols or which homes to inspect — trained on yesterday&#39;s data, they faithfully reproduce yesterday&#39;s unfairness, aimed at the same communities that have always been over-policed and under-served. The bias isn&#39;t in anyone&#39;s intentions. It&#39;s baked into the history the model learned from, which is exactly why it keeps happening.&lt;/p&gt;
&lt;p&gt;The one that makes me angriest is exclusion dressed as progress. A service that &lt;em&gt;requires&lt;/em&gt; an app, or a digital ID, or being online — that&#39;s a service that has decided the elderly, the homeless, and the newly arrived don&#39;t count. A city that locks people out of the basics to look modern has failed at the only thing a city is for.&lt;/p&gt;
&lt;h2 id=&#34;what-id-write-into-the-contract-in-ink&#34;&gt;What I&#39;d write into the contract, in ink&lt;/h2&gt;
&lt;p&gt;Some places have already worked out the right defaults, and I&#39;d shamelessly copy them. Amsterdam publishes a register of what each of its algorithms does and how you can challenge it. Estonia lets residents see who looked at their data and why. Copenhagen shares back what the city collects. Barcelona uses tech to &lt;em&gt;summarise&lt;/em&gt; what citizens say, not to replace asking them.&lt;/p&gt;
&lt;p&gt;Translate all that into plain contract language and it comes out as a short, stubborn list. Every system that makes decisions about residents is documented publicly. Collect the least data, keep it the shortest time, don&#39;t stitch datasets together on a whim. No facial recognition in public spaces — full stop, no asterisk. The city owns its own data. And you ask people before you build, not after, in a way that can actually change the answer.&lt;/p&gt;
&lt;p&gt;So here&#39;s the test I&#39;d apply to anything in that basket, on behalf of the man on the high street. Does it need to identify individuals to do its job? And if it goes wrong or gets breached, who pays, and can it be undone? The adaptive traffic lights pass both, easily. Public facial recognition fails both, badly. Most things sit in between — and that uncomfortable middle is precisely where someone has to fight for the right clauses.&lt;/p&gt;
&lt;p&gt;The best cities I&#39;ve seen aren&#39;t the ones bristling with the most sensors. They&#39;re the ones that decided who they were building for, and what they&#39;d refuse to build, before a salesperson decided it for them.&lt;/p&gt;]]></content:encoded>
    </item>
  </channel>
</rss>