{"site":"https://heliosmsp.io","customer":"Helios","accent":"#2c5ef0","posts":[{"slug":"msp-tooling-cost-percentage-revenue","title":"MSP Tooling Costs as a Percentage of MRR: How to Model It","meta_description":"How to model your MSP tooling cost percentage of revenue: build a tooling ledger, allocate costs per client and find the contracts quietly losing money.","published_at":"2026-08-12T21:13:50.103832","body_html":"<p>By the Helios team</p><p>Most MSP owners can quote their MRR to the pound and their tooling cost not at all. Ask what the RMM, PSA, EDR, backup and email security stack costs per managed endpoint per month and you get a shrug, because the answer is spread across six invoices, three billing units and one annual renewal nobody has looked at since it auto-renewed. Yet your MSP tooling cost percentage of revenue is the number that decides whether a new client at your standard rate makes money or quietly loses it. Here is a spreadsheet method that produces it, column by column, whether the machines belong to clients or to your own company.</p><h2>Start with a tooling ledger, not a feeling</h2><p>Open a sheet with five columns: tool, monthly cost, billing unit, quantity, and what it actually covers. Then list everything that touches service delivery: RMM, PSA, endpoint protection, backup, email security, documentation platform, remote access, password manager, phone system, the reporting add-on you bought in a weak moment.</p><p>Two disciplines matter here. First, normalise everything to a monthly figure: annual invoices divided by twelve, three-year commitments divided by thirty-six. Second, record the billing unit honestly, because it drives everything downstream. Most tools bill one of four ways: per endpoint, per technician, per client or tenant, or flat. If you are not sure which model you are on, that is itself a finding. Different vendors chose different units for reasons that suit them, not you, a point we have made at length when comparing <a href=\"https://helioseo.io/feed/helios/msp-pricing-models\">per-device, per-user and flat-fee pricing models</a>.</p><div class=\"callout\"><p><strong>Rule of thumb:</strong> convert every tool, whatever its billing unit, into cost per endpoint per month. It is the only unit that lets you compare tools with each other and costs with revenue.</p></div><h2>Allocate every cost to a client</h2><p>The ledger tells you what you spend. The allocation tells you who you spend it on. Three rules cover almost everything:</p><ul><li><strong>Per-endpoint tools allocate directly.</strong> If the EDR is £1.80 per endpoint and Client A has 60 endpoints, Client A carries £108. No judgement required.</li><li><strong>Per-technician and flat tools allocate by endpoint share.</strong> A PSA at £70 per technician for three technicians is £210 a month. Divide by your total endpoint count and multiply by each client's endpoints. It is imperfect, but it is consistent, and consistency is what makes the comparison between clients meaningful.</li><li><strong>Per-client tools allocate directly.</strong> A backup subscription tied to one client's server belongs on that client's row and nowhere else.</li></ul><p>The output is one row per client with a single figure: tooling cost per endpoint per month. Across a typical small MSP stack this lands somewhere in single-digit pounds, but the exact figure matters less than the spread. If one client costs twice the average, you want to know why before their renewal, not after.</p><h2>Gross margin per endpoint: the arithmetic</h2><p>Tooling is only half the cost of serving a client. The other half is labour, and your PSA already holds it: logged time per client per month. Take each technician's fully loaded monthly cost, work out a cost per logged hour, and multiply by the hours each client consumed. Then:</p><p><strong>Gross margin per endpoint = revenue per endpoint, minus tooling per endpoint, minus labour per endpoint.</strong></p><table><tr><th>Line</th><th>Client A (60 endpoints)</th><th>Client B (45 endpoints)</th></tr><tr><td>Revenue per endpoint</td><td>£28.00</td><td>£22.00</td></tr><tr><td>Tooling per endpoint</td><td>£5.20</td><td>£5.20</td></tr><tr><td>Labour per endpoint</td><td>£9.50</td><td>£19.80</td></tr><tr><td>Gross margin per endpoint</td><td>£13.30</td><td>£(3.00)</td></tr></table><p>Client B is paying you to lose money. Nothing on the invoice would tell you this, because the invoice shows revenue and revenue looks fine. Only the allocation shows it, and it usually shows it for a reason: legacy pricing that never rose, an estate full of ageing hardware, or a contact who treats the service desk as a colleague. Each has a different fix, and one of them is <a href=\"https://helioseo.io/feed/helios/msp-client-offboarding-checklist\">losing the client well</a>.</p><h2>MSP tooling cost as a percentage of revenue: the honest answer</h2><p>You will find people online claiming a correct benchmark for tooling as a percentage of MRR. Treat these numbers with suspicion. Community threads and peer surveys put the figure anywhere from mid single digits to around a fifth of revenue, and that spread is the real information: it depends entirely on your stack, your pricing and your endpoint density per client. An industry average applied to your business is a guess wearing a spreadsheet's clothes.</p><p>The percentages that are worth tracking are your own. Two of them: the whole-business figure, tracked quarter on quarter so you see drift when a vendor raises prices or a tool creeps in; and the per-client figure, because the variance between clients is where the unprofitable contracts hide.</p><h2>Model a price rise and a new hire before you commit</h2><p>Once the sheet exists, the interesting questions cost you nothing to answer.</p><p><strong>A price rise</strong> is simple: revenue per endpoint moves, tooling and labour do not, so the margin change drops straight through. A £2 per endpoint increase across 400 endpoints is £800 a month of pure margin, which is a useful number to hold in mind when you are nervous about the conversation.</p><p><strong>A new hire</strong> is where billing units bite. Every per-technician tool steps up on day one, before the hire has touched a ticket or brought in an endpoint. On a per-technician RMM and PSA, hiring your fourth technician can add a few hundred pounds a month of tooling cost with zero new revenue, which means your tooling cost per endpoint rises across every client simultaneously. Per-endpoint tools do the mirror image: they punish growth instead, taxing every new device you win. Run both scenarios in the sheet and you will see which billing units your business can afford to grow into. This is also why <a href=\"https://helioseo.io/feed/helios/how-to-choose-an-rmm\">choosing an RMM by feature checklist</a> misses the point: the pricing model shapes your margins for years.</p><div class=\"callout\"><p><strong>The renewal test:</strong> before signing or renewing any tool, add it to the ledger, reallocate, and look at what it does to gross margin per endpoint on your three thinnest clients. If any of them goes negative, the tool is not cheap, whatever the invoice says.</p></div><h2>Three ways this model lies to you</h2><ul><li><strong>Forgotten annual renewals.</strong> The tool billed once a year in March is invisible for eleven months. Skip the normalisation step and your percentage looks better than it is until the renewal lands.</li><li><strong>Averaged labour.</strong> Spreading labour evenly across clients instead of using logged time hides exactly the clients you built the sheet to find. If time logging is patchy, fixing that comes first.</li><li><strong>Treating the percentage as a target to minimise.</strong> The cheapest stack is rarely the cheapest service. A tool that saves each technician an hour a week is worth far more than its line in the ledger, and a bargain tool that generates noise costs you in labour instead. Minimise cost per endpoint served properly, not cost per endpoint on paper.</li></ul><h2>Where this fits with Helios</h2><p>Most of this exercise is discipline and a spreadsheet, not tooling. Where Helios changes the model is the billing unit: RMM, PSA, patching, remote access, service desk and client portal are one flat monthly line at £99, £199 or £399, banded by device count rather than priced per technician or per endpoint. That makes the allocation trivial and, more usefully, it means hiring a technician or winning fifty endpoints does not move your tooling cost at all. We should be plain about the limits: Helios does not yet do recurring contract billing or rate cards, so the revenue side of your sheet still comes from your accounting system, with time tracking exported to QuickBooks for invoicing.</p><p>Helios is a single RMM and PSA platform for small MSPs and internal IT teams, with flat monthly pricing and every feature on every plan. There is a 14-day trial with no card and no feature gating. Start free at <a href=\"https://heliosmsp.io\">heliosmsp.io</a>.</p>"},{"slug":"ai-automation-guardrails-it","title":"AI Guardrails for IT Automation: Approvals, Scopes and Audit Trails That Keep You Safe","meta_description":"AI automation guardrails for IT: approval modes, allow-lists, blast-radius limits and audit trails, plus a written policy you can show a nervous client.","published_at":"2026-08-12T20:57:28.563009","body_html":"<p>By the Helios team</p><p>The problem with AI automation is not that it makes mistakes. Technicians make mistakes too. The problem is that it makes them at machine speed, across every device it can reach, at three in the morning when nobody is watching. That is why AI automation guardrails in IT are not an optional extra for the cautious: they are the whole difference between an assistant and a liability. This piece sets out the guardrails that matter, per-client approval modes, allow-lists, blast-radius limits, dry runs, verification and immutable audit logs, and ends with a written policy you can hand to a nervous client.</p><h2>What AI automation guardrails in IT actually are</h2><p>Strip away the vendor language and a guardrail is one of three things: a limit on <strong>what</strong> an automation may do, a limit on <strong>where</strong> it may do it, and a record of <strong>what it did</strong> that nobody can quietly edit afterwards. Any system missing one of the three is not guarded. It is merely supervised, and only for as long as someone is looking.</p><p>The temptation to skip guardrails is real, because the queue is real. Alert volumes push small teams towards automation long before they have thought about failure modes, and we have written before about <a href=\"https://helioseo.io/feed/helios/msp-alert-fatigue\">cutting alert noise without missing real incidents</a>. Automation is the right answer to that pressure. Unbounded automation is not.</p><h2>Per-client approval modes: not everyone deserves the same autonomy</h2><p>The first mistake teams make is treating autonomy as a global switch. It is a per-client setting, whether the machines belong to clients or to your own company's departments. A sensible model has three modes:</p><ul><li><strong>Suggest only.</strong> The AI investigates, writes up its diagnosis and proposed fix, and stops. A human runs it or bins it. This is the default for every new client and every new automation, no exceptions.</li><li><strong>Approve to run.</strong> The AI queues the fix and a technician approves it with one click. The human is still the gate, but the investigation and scripting work is already done. Most clients should live here for months.</li><li><strong>Autonomous within scope.</strong> The AI acts without asking, but only for actions on the allow-list, only within the blast-radius limit, and only with verification and logging. This mode is earned, not enabled.</li></ul><blockquote><p>Autonomy is not a setting you switch on. It is a privilege an automation earns, one verified fix at a time.</p></blockquote><p>The per-client part matters commercially as well as technically. A law firm with regulatory exposure and a five-person joinery with three laptops should not share an approval posture, and being able to show a client their specific settings is worth more in a renewal conversation than any amount of reassurance.</p><h2>Allow-lists, not block-lists</h2><p>A block-list says \"do anything except these things\" and fails the moment something you did not anticipate happens, which is the only kind of thing that ever happens. An allow-list says \"do nothing except these things\" and fails safe. Start with a short list of boring, reversible actions: restart a named service, clear a temp directory, re-register a Windows Update component, restart the print spooler. Add to it deliberately, one action at a time, after each has run cleanly under approval mode for a while.</p><p><strong>The reversibility test:</strong> before any action goes on the allow-list, ask whether a technician could fully undo it in under five minutes with the information in the log. Restarting a service passes. Deleting files matching a pattern does not, because you cannot un-delete your way out of a bad pattern.</p><h2>Blast radius: the limit that saves you at 3am</h2><p>Even an approved, allow-listed action becomes dangerous at scale. A fix that is correct for one machine can be wrong for a hundred, because the diagnosis was pattern-matched rather than reasoned, or because a hundred simultaneous restarts is itself an incident. So cap the radius: per device, per client, per hour.</p><div class=\"callout\"><p><strong>Rule of thumb:</strong> never let an unattended automation touch more devices in an hour than your team could manually put right in a day. For a two-technician shop, that is a handful of machines, not a fleet.</p></div><p>Blast-radius limits also interact with patching, where the temptation to go wide and fast is strongest. Rings and caps feel slow, but as we argued in <a href=\"https://helioseo.io/feed/helios/how-fast-should-you-patch\">how fast should you patch</a>, the answer is to make each ring short, not to remove the rings.</p><h2>Dry runs and verification: check before, check after</h2><p>Two habits separate automation you can defend from automation you have to apologise for.</p><ul><li><strong>Dry runs.</strong> Any new script or fix should be able to report what it would do without doing it. If the dry run output surprises you, the live run would have surprised you more. Skip this and the first real test of your automation happens in production, on a client's machine, with their name on the incident.</li><li><strong>Verification after the fix.</strong> An automation that restarts a service and walks away has not fixed anything, it has performed a ritual. The fix is complete when the automation re-checks the original symptom: the service is running, the disk has space, the alert has cleared. If verification fails, the correct behaviour is to stop and escalate to a human, not to try something else. Retrying with escalating creativity is how a stuck update becomes a broken machine.</li></ul><h2>Audit logs that cannot be edited</h2><p>Every automated action needs a record of what triggered it, what the AI concluded, what it ran, on which device, in whose environment, under which approval mode, and what the verification showed. And the record must be immutable, because an audit trail you can edit is a diary, not evidence. When a client asks \"what did your robot do to my server last Tuesday\", the answer should be a timestamped log you can export, not a reconstruction from memory. This is the same discipline we argue for in the <a href=\"https://helioseo.io/feed/helios/msp-security-checklist\">MSP security checklist</a>: hold your own tooling to the standard you would demand from a supplier.</p><h2>Safe unattended versus never unattended</h2><table><tr><th>Action</th><th>Unattended?</th><th>Why</th></tr><tr><td>Restart a named service on one device</td><td>Yes</td><td>Reversible in seconds, narrow scope, easy to verify</td></tr><tr><td>Clear temp files from known safe paths</td><td>Yes</td><td>Bounded, well-understood, low blast radius</td></tr><tr><td>Apply approved patches within a ring</td><td>Yes, with caps</td><td>Pre-approved content, ringed rollout, verification built in</td></tr><tr><td>Delete files matching a pattern</td><td>Never</td><td>Fails the reversibility test; a bad pattern is unrecoverable</td></tr><tr><td>Change firewall, DNS or identity settings</td><td>Never</td><td>Can sever your own access and lock out users at once</td></tr><tr><td>Disable or reset user accounts</td><td>Never</td><td>Business impact is human, not technical; needs human judgement</td></tr><tr><td>Anything touching backup jobs or retention</td><td>Never</td><td>Backups are the safety net; the net does not get automated holes</td></tr></table><h2>The written policy your nervous client can read</h2><p>Put all of this on one page, per client, in plain language:</p><ol><li><strong>Approval mode.</strong> Which of the three modes applies to this client, and who can change it.</li><li><strong>The allow-list.</strong> The specific actions permitted unattended, named, not described in categories.</li><li><strong>Blast-radius limits.</strong> Maximum devices per action, per hour, per client.</li><li><strong>Verification and escalation.</strong> Every unattended fix is re-checked; every failed verification goes to a human within a stated time.</li><li><strong>The audit trail.</strong> Every action is logged immutably and the client can request the log at any time.</li><li><strong>The never list.</strong> The actions that will not run unattended under any mode, in writing.</li></ol><p>A nervous client is not asking you to promise the AI will never be wrong. They are asking you to show that when it is wrong, it will be wrong in a small, reversible, fully documented way. This page answers that.</p><h2>Where this fits with Helios</h2><p>Most of this article is discipline, and no platform can supply discipline. What Helios supplies is the machinery: our AI agent, Helio, investigates alerts and writes fixes, but every action runs under per-client approval modes, with verification after each fix and a full log of what it concluded and what it ran. Suggest-only is the default for every new client, and autonomy is something you grant scope by scope, not a global switch. We built it this way because we run an MSP on it ourselves and we are exactly as nervous as your clients are.</p><p>Helios is a flat-fee RMM and PSA with an AI agent built in, from £99 a month with every feature on every plan. 14-day trial, no card, no feature gating. <a href=\"https://heliosmsp.io\">Start free</a>.</p>"},{"slug":"msp-security-checklist","title":"MSP security checklist: harden your own house first","meta_description":"A practical MSP security checklist: six controls that decide whether one phished technician credential becomes fifty breached client estates.","published_at":"2026-08-12T00:00:00","body_html":"<p class=\"lede\">Every MSP sells security. Far fewer apply the same rigour to their own house, and attackers know it. Compromise one company and you get one estate; compromise an MSP and you get every estate it manages, delivered through tooling those clients are contractually required to trust. This MSP security checklist covers the six controls that decide whether one phished technician credential stays an incident or becomes a supply chain breach with your name on it.</p>\n\n    <p>Each item below gets a short explanation of why it matters and what tends to happen when it is skipped. Work through them in order: the first three shrink the blast radius of a compromise, the last three make sure a compromise stays survivable.</p>\n\n    <h2>1. Put phishing-resistant MFA on everything that can touch a client estate</h2>\n    <p>Your RMM, your PSA, your password manager, your documentation platform and your Microsoft partner portal are, taken together, a master key to every client you serve. ConnectWise's 2026 MSP threat report puts identity abuse at the centre of MSP risk for a reason: nobody needs to exploit software when a convincing login page and a tired technician will do. Ordinary push-based MFA is no longer enough for these accounts, because push fatigue and real-time phishing proxies defeat it routinely. For anything with cross-client reach, that means passkeys or hardware security keys, plus conditional access rules that refuse logins from unknown devices. Start by listing which accounts actually have that reach; most MSPs find the list is longer than they expected once the documentation platform, the backup consoles and the odd forgotten vendor portal are counted.</p>\n    <p>Skip this and the failure mode is brutally simple: one technician approves one prompt at the end of a long day, and an attacker is inside a console with interactive access to every device you manage. Most of the MSP breaches that make the news start exactly here, not with a zero-day.</p>\n\n    <h2>2. Give technicians per-client identities, not one master key</h2>\n    <p>Shared admin accounts and standing global access are convenient, and that convenience is precisely what makes them dangerous. Every technician should reach each client tenant through a separate, named identity with the least privilege the job requires, elevated only when needed and logged when used. Microsoft partners should have finished the move from legacy delegated admin (DAP) to granular delegated admin (GDAP) long ago; if that migration is still on your backlog, it belongs at the top. The same <a class=\"inline\" href=\"/blog/microsoft-365-security-checklist\">Microsoft 365 hardening baseline</a> you apply to client tenants applies doubly to your own. Then review the access map quarterly, because privilege only ever accumulates: technicians change teams, temporary elevations become permanent, and a year later nobody can say who can reach what.</p>\n    <p>Skipped, this control turns any single compromise into a full portfolio event. Flat access means lateral movement between clients costs an attacker nothing, and it also means your incident report has to say \"we cannot rule out access to any customer\", which is the sentence that ends client relationships.</p>\n\n    <h2>3. Inventory every remote access path, then close the ones you cannot defend</h2>\n    <p>Remote access tooling is now a primary attack vector in its own right. Huntress reported a 277% jump in RMM abuse during 2025, accounting for roughly a quarter of the incidents it investigated, and CISA spent June 2025 warning about ransomware crews exploiting unpatched SimpleHelp instances. Attackers either hijack the legitimate agent you deployed or quietly install one of their own, because a signed, well-known remote access tool walks straight past most defences. The checklist item is threefold: sanction exactly one remote access product, patch it within days of a release rather than months, and alert the moment any other remote access binary appears on a device you manage.</p>\n    <div class=\"callout\"><p><strong>The second-RMM blind spot:</strong> an attacker does not need to compromise your RMM if they can install their own next to it. A spare ScreenConnect or AnyDesk instance sitting on a client server is one of the most common persistence mechanisms seen in real intrusions, and it is invisible unless something is explicitly watching for tools you did not deploy.</p></div>\n    <p>Without this control, an intruder's access outlives every password reset you do, because the backdoor is not a credential. It is a legitimate product, running as a service, that nobody on your team remembers installing.</p>\n\n    <h2>4. Run your own estate as client zero</h2>\n    <p>The cobbler's children go barefoot in almost every MSP: internal laptops miss patch windows that client machines never would, the office server has no backup alerting, and nobody monitors the technicians' endpoints with the rigour applied to a client's finance team. That inversion is exactly backwards. A technician's laptop holds live sessions and cached tokens for your RMM, your PSA and a dozen client tenants, which makes it more valuable to an attacker than any device you manage for anyone else. The same goes for your documentation platform: a well-kept IT documentation store is a beautifully organised credential database from an intruder's point of view, and it deserves the same protection as the systems it describes.</p>\n    <p>The fix is procedural rather than technical: enrol your own company as a tenant in your own stack, with the same patch SLAs, the same endpoint protection checks and the same alerting as your best-managed client. When this is skipped, the intrusion route is depressingly consistent: the attacker never touches your hardened RMM platform at all. They compromise the unpatched, unmonitored laptop that is already logged into it.</p>\n\n    <h2>5. Offboard leavers with the urgency of a zero-day</h2>\n    <p>When a technician resigns, their knowledge and credentials span every client on your books, and every hour of delay is exposure multiplied by your whole customer list. That demands a living inventory of what each role can touch: which tenants, which consoles, which shared secrets, which client-site accounts. On their last day, named accounts get disabled everywhere, and every shared secret they knew gets rotated, from documentation-platform entries to client local admin passwords. Our <a class=\"inline\" href=\"/blog/it-offboarding-checklist\">IT offboarding checklist</a> covers the general case; for an MSP, multiply every step by the number of clients involved.</p>\n    <p>Skip it and you are gambling on goodwill. Most leavers are honest, but \"former employee of the IT provider\" appears in breach post-mortems far too often to be a hypothetical, and even an honest leaver's stale account is a ready-made target for whoever phishes it next.</p>\n\n    <h2>6. Rehearse the day your own tooling is the breach</h2>\n    <p>Every MSP has an incident response plan for clients. Very few have one for the scenario where the incident is them: ransomware deploying through their own RMM, to every client at once. Write that plan while it is still hypothetical. It needs a tested way to mass-disable or isolate your agents, break-glass credentials stored offline where a compromised tenant cannot reach them, an out-of-band contact list for clients that does not live in the Microsoft 365 tenant you may have just lost, and an honest decision, made in advance, about who tells clients what and when. Then tabletop it twice a year, and check your own cyber insurance actually covers incidents that propagate to customers, because a policy written for a normal business often excludes exactly the third-party liability an MSP breach creates.</p>\n    <p>Without a rehearsed plan, the first hours of a real event are spent arguing about basics while the attacker uses your deployment tooling at full speed. Those hours are the difference between isolating three clients and explaining yourself to fifty.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios cannot run your tabletop exercise for you, but it makes \"client zero\" practical: your own company sits in the platform as a tenant like any other, so your laptops and servers get the same patching, endpoint protection checks and backup monitoring as every client estate. Protection-gap detection flags devices where security tooling is missing or silent, on your estate as much as theirs, and Helio's device investigations give you a fast answer when something on a managed machine does not look like something you deployed. An MSP security checklist only works if checking it is cheap enough to do continuously, and that is the part a platform should carry.</p>"},{"slug":"3-2-1-backup-rule","title":"The 3-2-1 backup rule: still right, no longer enough","meta_description":"The 3-2-1 backup rule is the most repeated advice in IT, and it is still sound. Here are the three modern failures it quietly says nothing about.","published_at":"2026-08-11T00:00:00","body_html":"<p class=\"lede\">Ask anyone in IT how to do backups properly and you will hear the same answer: three copies of your data, on two different types of media, with one copy offsite. The 3-2-1 backup rule has been standard advice for two decades, and whether you run your own company's estate or look after backups for clients, you have probably repeated it yourself. It is still good advice. It has also quietly stopped being sufficient, and the gap between those two facts is where businesses get hurt.</p>\n\n    <h2>What the 3-2-1 backup rule gets right</h2>\n    <p>The rule endures because it encodes one genuinely important idea: no two copies of your data should be able to fail for the same reason. Disks die, so keep more than one copy. A whole class of media can fail or be corrupted together, so use two kinds. A building can flood, burn or be burgled, so keep a copy somewhere else.</p>\n    <p>Against the failures it was designed for, accidents, hardware faults and local disasters, 3-2-1 still works. It is short enough to remember, simple enough to audit, and it has saved an enormous amount of data. Nothing below argues you should abandon it. The problem is what it never claimed to cover, and what people assume it covers anyway.</p>\n\n    <h2>Myth one: an offsite copy is a safe copy</h2>\n    <p>The rule was written for a world where data loss was accidental. Ransomware is not an accident, it is an adversary, and modern operators go for the backups first: they delete shadow copies, sign in to the backup console with stolen admin credentials, and purge cloud repositories before they encrypt anything. If your offsite copy can be reached, changed or deleted using the same credentials that run the rest of your estate, it is not a separate copy in any meaningful sense. It shares a fate with the primary.</p>\n    <p>This is why the industry has drifted towards <strong>3-2-1-1-0</strong>: the extra 1 is a copy that is immutable or genuinely offline, something a compromised administrator account cannot touch. You do not need the numerology, but you do need the property. At least one copy should survive an attacker who owns your credentials.</p>\n    <div class=\"callout\"><p><strong>Sync is not backup.</strong> OneDrive, Dropbox and Google Drive replicate whatever happens to your files, including deletion and encryption, usually within minutes. A synced copy fails the 3-2-1 test in spirit even when it appears to pass it on paper, because it is designed to mirror damage faithfully.</p></div>\n\n    <h2>Myth two: a completed job is a usable backup</h2>\n    <p>3-2-1 counts copies. It says nothing about whether any of them can actually be restored, and that is the failure mode we see most often in real estates: the job that has been silently failing for six weeks, the chain that broke after a server was renamed, the backup that completes nightly but has the wrong folders in scope, the restore that technically works but takes four days when the business assumed four hours.</p>\n    <p>The 0 in 3-2-1-1-0 stands for zero errors after verification, and it is the least glamorous, most valuable part of the whole formula. Verification means restoring real data on a schedule and timing it, not reading a dashboard. We have written before about <a class=\"inline\" href=\"/blog/backup-monitoring-for-msps\">why a green tick is not a restore</a>, and the short version holds: a backup you have never restored from is a hypothesis, not a control.</p>\n\n    <h2>Myth three: following the rule means everything is covered</h2>\n    <p>The rule protects the data inside your backup jobs. It is silent about the data that never made it into one. Three gaps come up constantly:</p>\n    <ul>\n      <li><strong>Microsoft 365 and other SaaS data.</strong> Microsoft keeps the service running; under the shared responsibility model, long-term recovery of your data is your job. Retention policies are not backup, and plenty of estates with immaculate server backups have nothing behind Exchange Online or SharePoint. Our <a class=\"inline\" href=\"/blog/microsoft-365-security-checklist\">Microsoft 365 security checklist</a> pairs well with fixing this.</li>\n      <li><strong>Laptops.</strong> Users keep real work on local disks whatever the policy says, and most endpoint fleets are backed by nothing but hope and OneDrive.</li>\n      <li><strong>New systems.</strong> The server stood up in March that nobody added to a job. Scope rot is gradual and invisible.</li>\n    </ul>\n    <p>The fix is to start from inventory, not from the backup tool: list every system and dataset the business would miss, then check each one against 3-2-1. Auditing only the things that are already in a job tells you nothing about the things that are not.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios watches Veeam and Acronis jobs across every site you look after and treats silence as a finding: failed jobs, stale jobs and jobs that have never run all surface in one view, next to the asset inventory, so a device with no backup at all is visible rather than simply absent. That is the unglamorous half of 3-2-1-1-0, the verification and the coverage, handled by the same agent that already tracks patching and protection state.</p>"},{"slug":"business-email-compromise","title":"Business email compromise: anatomy of an attack that almost worked","meta_description":"Business email compromise rarely starts with the fake invoice. We trace one attack backwards, from payment to first sign-in, and the changes that stop the next.","published_at":"2026-08-10T00:00:00","body_html":"<p class=\"lede\">The finance manager finds out on a Tuesday, when a supplier rings to ask why an invoice is six weeks overdue. It was paid, finance insists. It was, in fact, paid to someone else. That phone call is how most business email compromise comes to light: not an alarm, but an awkward conversation long after the money moved. What follows is a post-mortem of a typical incident, a composite of cases we have seen, traced backwards from the payment to the root cause, then forward to the changes that stop the next one. It applies whether you look after one company's estate or thirty clients' worth.</p>\n\n    <h2>The phone call that starts the post-mortem</h2>\n    <p>The paper trail looks clean at first. Finance received the supplier's invoice as usual, on the usual thread, quoting the right PO number. A few days later came a follow-up from the same thread: the supplier had \"changed banks\" and would finance kindly use the new account details. The email read naturally, referenced the right people, and arrived at a believable moment. Finance updated the record and paid. Nothing about it looked like fraud, because almost none of it was fake: the attacker was inside a real mailbox, replying inside a real conversation.</p>\n    <p>This is worth stating plainly, because the popular image of email fraud is still a badly spelled message from a stranger. Business email compromise is the opposite. In the FBI's 2025 internet crime figures it accounted for just over <strong>$3 billion in reported losses across roughly 25,000 complaints</strong>, an average of more than $120,000 per incident, and 86% of it moved by wire or ACH transfer, which means it was largely gone before anyone noticed. Those are only the reported cases. The typical victim is not a careless person; it is an ordinary payment process that trusted an email it had every apparent reason to trust.</p>\n\n    <h2>Working backwards: the rule hiding in RSS Feeds</h2>\n    <p>The first real finding in the investigation is not the fraudulent email. It is an inbox rule in the finance manager's mailbox, created weeks before the payment, that nobody made. The rule is simple: any incoming message containing words like \"invoice\", \"payment\" or \"bank details\", or anything from the supplier's domain, gets marked as read and moved to the RSS Subscriptions folder. A folder that exists in every Outlook mailbox and that no human being has ever opened on purpose.</p>\n    <p>The rule is the attack. It let the intruder lift a genuine invoice out of the conversation before finance saw it, send their doctored follow-up, and then quietly swallow the supplier's replies, including the one that said \"we have not changed banks, please ignore that email\". It also swallowed anything that might have raised suspicion, like security notifications. And because inbox rules live in the mailbox rather than the session, the rule kept working even after the user later changed their password for unrelated reasons. Rules like this are not exotic: incident responders consistently find that a large share of compromised Microsoft 365 accounts, by some counts around 40%, get a malicious mailbox rule within minutes of the break-in. In one documented case the gap between compromise and rule creation was eight seconds. It is the first thing the playbook does.</p>\n\n    <h2>Further back: the sign-in that passed MFA</h2>\n    <p>So who created the rule? The audit log answers that too: a sign-in from an unfamiliar network, three weeks before the payment, at 2:41 in the morning local time. Here is the uncomfortable part. That sign-in satisfied multi-factor authentication. The organisation had MFA enabled, the user had it registered, and the attacker walked through it anyway.</p>\n    <p>The mechanism is adversary-in-the-middle phishing, and it is now the standard entry route. The user had received a convincing email a day earlier and clicked through to what looked exactly like the Microsoft sign-in page. It was a proxy: every keystroke was passed through to the real Microsoft login, including the genuine MFA challenge, which the user completed because from their side nothing was wrong. What the attacker kept was not the password but the <strong>session token</strong> issued at the end, the small credential that tells Microsoft 365 \"this browser has already authenticated\". Replay that token from anywhere and you are the user, with no password prompt and no MFA prompt, until the token expires or is revoked.</p>\n\n    <h2>The root cause was not the click</h2>\n    <p>At this point in most internal post-mortems, someone writes \"user clicked a phishing link\" in the root cause box and schedules more security awareness training. That is the wrong lesson, and it is worth being precise about why. The click was the trigger. The root cause was a set of conditions that turned one click into a five-figure loss:</p>\n    <ul>\n      <li><strong>A stolen session worked from anywhere.</strong> Nothing checked whether the token was being replayed from an unknown device or an unexpected country.</li>\n      <li><strong>The persistence step was invisible.</strong> A new inbox rule in a finance mailbox is one of the loudest signals a tenant produces, and nothing was listening for it.</li>\n      <li><strong>The payment control lived inside email.</strong> The process for changing a supplier's bank details was \"receive an email asking us to\".</li>\n    </ul>\n    <p>Every estate has users who will eventually click; the click rate never reaches zero, however much training you buy, and <a class=\"inline\" href=\"/blog/do-phishing-simulations-work\">the evidence on phishing simulations</a> says they barely move it. A control that only works if no human ever clicks a link is not a control. The honest root cause statement is: our security model assumed authentication happens once, at sign-in, and everything afterwards is trusted. The attacker simply operated in the \"afterwards\".</p>\n\n    <h2>Why three weeks passed unnoticed</h2>\n    <p>The most painful line in the timeline is the gap. Between the 2:41am sign-in and the supplier's phone call, twenty-three days passed. In that time the attacker signed in repeatedly, searched the mailbox for \"invoice\" and \"payment run\", read the finance calendar, created the rule, and sent the bank-change email. Every one of those actions was logged. Not one log was read.</p>\n    <p>This is the quiet failure underneath most business email compromise: the evidence was available the whole time, and availability is not the same as attention. Small and mid-sized organisations, and the lean teams and MSPs who support them, tend to assume mailbox compromise is something that happens to bigger targets. It is not. Attackers do not shortlist victims; they automate, and a 40-person firm's finance mailbox pays out just as well as anyone's. A server going offline pages somebody at 3am. A foreign sign-in to the finance manager's mailbox, in most estates, pages nobody at all. That asymmetry, loud infrastructure and silent identity, is exactly what this class of attack exploits.</p>\n\n    <h2>The changes that prevent the next business email compromise</h2>\n    <p>The fixes fall out of the timeline in reverse order, and none of them is training. In rough order of value:</p>\n    <ol>\n      <li><strong>Take bank details out of email entirely.</strong> Any change to payment details gets verified by a phone call to a number already on file, never one from the email. This is a policy, it costs nothing, and it would have stopped this incident on its own even with everything else failing.</li>\n      <li><strong>Alert on the persistence moves.</strong> New inbox rules that move or delete mail, new forwarding addresses, new OAuth application consents, new MFA methods registered. These are rare, high-signal events. In this incident, a single alert on rule creation would have cut the dwell time from three weeks to one morning.</li>\n      <li><strong>Make stolen tokens worth less.</strong> Conditional access policies that require a compliant or known device for mail access, at minimum for finance and admin roles, mean a replayed session from an attacker's server fails even with a valid token. Block legacy authentication protocols while you are there; they skip MFA entirely.</li>\n      <li><strong>Move the riskiest people to phishing-resistant MFA.</strong> Passkeys and FIDO2 keys bind authentication to the real site, so a proxy page has nothing to steal. Rolling this out to finance, executives and IT admins covers most of the payout targets with a fraction of the effort of a full rollout.</li>\n      <li><strong>Rehearse revocation.</strong> Know, before you need it, how to revoke a user's sessions in under five minutes. Resetting the password alone does not end a token-based intrusion, and it does not remove the rule.</li>\n    </ol>\n    <p>If you have not hardened the tenant recently, our <a class=\"inline\" href=\"/blog/microsoft-365-security-checklist\">Microsoft 365 security checklist</a> walks through the ten settings that close most of this off, including the conditional access and legacy authentication pieces above.</p>\n\n    <h2>If you find one today: the first hour</h2>\n    <p>Post-mortems are for the next incident. If you are reading this because you have just found the rule in someone's RSS folder, order matters:</p>\n    <ol>\n      <li><strong>Revoke sessions first</strong>, then reset the password. The reset alone leaves live tokens working.</li>\n      <li><strong>Delete every inbox rule and forwarding address you did not create</strong>, and check the account's registered MFA methods and recently consented apps for anything the attacker added to keep a way back in.</li>\n      <li><strong>Call the bank now</strong>, not after the investigation. Recalling a fraudulent transfer is measured in hours; after a day or two the money has usually been layered out of reach.</li>\n      <li><strong>Assume lateral spread.</strong> Check other mailboxes for the same rule patterns and sign-in sources; attackers who land in one mailbox routinely phish colleagues from it, because internal mail is trusted mail.</li>\n      <li><strong>Report and record.</strong> Action Fraud in the UK, IC3 in the US, and your insurer. Preserve the audit logs before retention quietly ages them out.</li>\n    </ol>\n\n    <div class=\"callout\"><p><strong>One test worth running this week:</strong> pick a finance mailbox and create a harmless inbox rule that moves a keyword to an obscure folder. Time how long it takes anyone, or anything, to notice. That number is your current business email compromise dwell time. For most estates the honest answer is \"indefinitely\", and now you know which control to build first.</p></div>\n\n    <h2>Where this fits with Helios</h2>\n    <p>The pattern in this post-mortem, evidence that existed but was never looked at, is the problem Helios is built around. Helios connects to Microsoft 365 so identity sits in the same view as devices, patching, antivirus and backups, and Helio, the AI layer, treats an anomalous sign-in or a suspicious new mailbox rule the way it treats a failing disk: as something to investigate and surface now, not a log line to discover three weeks later. Whether you run IT for one business or for many, the goal is the same, to make the silent parts of the estate as loud as the servers.</p>"},{"slug":"how-to-choose-an-rmm","title":"How to choose an RMM: stop scoring feature lists","meta_description":"Every RMM ticks every feature box, so comparison matrices cannot pick a winner. How to choose an RMM on what actually differs: failure behaviour, defaults and the exit.","published_at":"2026-08-09T00:00:00","body_html":"<p class=\"lede\">Search for how to choose an RMM and you will find the same article twenty times: a checklist of features to look for, every one of which every serious product already has. This piece makes a different claim. Feature comparison can no longer pick an RMM, because the features have converged, and the differences that will actually cost you money live in places a comparison matrix never looks. That holds whether you are an MSP responsible for thirty client estates or an internal IT team buying proper tooling for the company you work for.</p>\n\n    <h2>The feature matrix stopped working years ago</h2>\n    <p>Ten years ago the matrix earned its keep. Some RMMs had patch management and some did not. Some could manage Macs, some could not. Remote access was a differentiator rather than an assumption. You could line up four products, tick boxes, and the winner genuinely was the better tool for you.</p>\n    <p>That world is gone. Monitoring, patching, scripting, remote access, software deployment, antivirus visibility, reporting, an API: every credible product ticks every row. The market converged for a boring reason, which is that a feature matrix is also the vendor's roadmap. Whenever a competitor's tick became a reason to lose a deal, the gap got filled, sometimes properly and sometimes with the minimum implementation that lets sales say yes.</p>\n    <p>So when you score five products against forty criteria and they finish within a few points of each other, that is not evidence they are interchangeable. It is evidence your instrument cannot measure the difference. The differences are still there, and some of them are enormous. They are just not feature-shaped.</p>\n\n    <h2>But surely the features do differ</h2>\n    <p>The obvious objection first: of course a tick is not a tick. One product's patch engine is superb and another's technically exists. True, and it proves the point rather than refuting it. If the same row on the matrix can hide both a superb implementation and a token one, the matrix is not carrying the information you need. The question is never \"does it have patch management\", it is \"what happens on the machine where patching goes wrong\", and no comparison table has a column for that.</p>\n    <p>The demo cannot rescue you either, because a demo is the vendor's best case by construction: a clean tenant, a handful of healthy devices, a rehearsed path through the happy flows. Every product looks excellent under those conditions, which means the demo, like the matrix, fails to discriminate. You are not buying the product's behaviour on its best day. You are buying its behaviour on your worst one, at 2am, on a laptop that has not checked in since Tuesday. The rest of this article is about the five places that behaviour actually varies: the agent, the defaults, the automation model, the pricing shape, and the exit.</p>\n\n    <h2>What actually differs: how the agent fails</h2>\n    <p>An RMM is an agent with a console attached, not the other way round. The console is what you are shown in the sales cycle; the agent is what you live with. And agents differ most not in what they collect but in how they behave when something goes wrong on the endpoint.</p>\n    <p>The failure mode that should frighten you is the quiet one. An agent that crashes leaves a trail. An agent that keeps checking in while silently failing to do part of its job leaves a green dashboard and a false belief. We have written before about <a class=\"inline\" href=\"/blog/third-party-patching-for-msps\">a patching integration that reports \"no updates found\" when it actually means \"could not look\"</a>, and that pattern generalises across the whole category: the worst agent bugs manifest as reassurance.</p>\n    <p>So interrogate the failure behaviour directly. Does the platform distinguish \"checked and found nothing\" from \"failed to check\"? What does it do with a device that is online but has stopped reporting one data type? How does the agent recover from a corrupted install, and how would you ever know it needed to? Ask each vendor these questions and you will learn more from the quality of the answers than from any datasheet. A vendor who can describe their agent's failure modes in detail has met them and fixed them. A vendor who insists the agent just works has customers doing their QA.</p>\n\n    <h2>The defaults are the product you really bought</h2>\n    <p>In principle every RMM is infinitely configurable, and vendors lean on this: whatever you dislike, you can change. In practice, teams run far closer to the defaults than anyone admits, because tuning monitoring policies is exactly the kind of important, non-urgent work that loses to a full ticket queue every single day. Two years in, most estates are running a lightly edited version of whatever the product shipped with.</p>\n    <p>That makes the out-of-box experience a preview of your long-term reality, and nowhere more than alerting. An RMM that installs on a hundred devices and produces eight hundred alerts in its first week has told you precisely what it thinks an alert is worth, and <a class=\"inline\" href=\"/blog/msp-alert-fatigue\">a team drowning in noise stops seeing real incidents</a> long before anyone gets around to fixing the thresholds. The volume you see in week one, untuned, is the honest signal. Judge it as such.</p>\n\n    <h2>The automation model is a worldview, not a checkbox</h2>\n    <p>Every product ticks \"automation\", and the tick hides the deepest difference in the category. One school gives you a script library and a scheduler: powerful, and quietly a commitment to author, test and maintain a codebase of remediations yourself, forever. That is a fine trade for a team with real engineering capacity, and a slow-motion failure for a team of two, because the library decays as Windows moves underneath it and the person who wrote it leaves.</p>\n    <p>The other school ships the remediation logic as part of the product, increasingly with AI attached. Here the risk inverts: instead of maintenance burden, opacity. If a platform advertises AI-driven anything, make it concrete. What exactly will it do without a human approving it? What is the worst action it can take unattended? Can you see, afterwards, precisely what it did and why? Confident, specific answers describe a real capability with real guardrails. Vague gestures at intelligence describe a demo.</p>\n    <p>Neither school is wrong, but they suit different teams, and this single axis should shape your shortlist more than thirty matrix rows put together. Buy the automation model you can realistically operate, not the one that impressed you.</p>\n\n    <h2>Pricing shape matters more than the price</h2>\n    <p>Two quotes for the same estate can differ by a third and still matter less than the shape of the number. Per-technician pricing punishes growing the team, so it quietly discourages hiring. Per-endpoint pricing scales smoothly but adds up fast across a large estate. Bundles bury the RMM inside a suite where the real per-unit cost is unknowable, which is precisely the point of them. We have covered this ground from the seller's side in <a class=\"inline\" href=\"/blog/msp-pricing-models\">our piece on MSP pricing models</a>, and the logic is symmetrical when you are the buyer: whatever the metric, you will optimise against it, so pick the metric you can live with optimising.</p>\n    <p>Then look at the term. A three-year commitment with an auto-renewal clause is not a discount, it is the vendor pricing in your future dissatisfaction. Among MSPs who switch RMMs, the most common trigger is not a missing feature but a contract that could not flex when the business changed. If the product is as good as the salesperson says, it can be that good on a rolling term.</p>\n\n    <h2>\"We can always migrate later\" is the costliest sentence in the deal</h2>\n    <p>The other standard objection to all this care: just pick something reasonable, and switch if it disappoints. Everyone in this industry knows how that ends, because everyone has met an MSP running an RMM they have hated for five years. Migration means touching every endpoint, rebuilding every policy and automation, retraining every technician, and re-learning every alert's meaning, all while the day job continues. The switching cost is so high that mediocre platforms retain customers for years on inertia alone. Vendors know this, which is why retention does not prove quality.</p>\n    <p>So evaluate the exit while you still have leverage, before signing. Can you export your data, all of it, in a usable format, without a professional services engagement? Can you mass-remove the agent cleanly? What does the contract say about your data after termination? A vendor confident in the product makes leaving easy, because they expect you to stay for better reasons. A hard exit is a confession.</p>\n\n    <h2>So how do you choose an RMM? Run the ugly trial</h2>\n    <p>The remaining shortcuts fail for the same underlying reason. \"Buy the market leader\" outsources the decision to other people's circumstances, and in a category reshaped by consolidation, the leader you buy is often an acquisition target whose roadmap and pricing you cannot predict. \"Buy the cheapest\" optimises the one number this article has argued matters least. Both are ways of avoiding the work of evaluating against your own estate, and your estate is the only benchmark that counts.</p>\n    <p>The method that does work is cheap and unglamorous. Shortlist two or three products on the axes above. Then trial each one on your fifty ugliest devices: the ancient tablet in the warehouse, the laptop that lives on hotel wifi, the server nobody dares reboot. Give it a month and score four things: how the agent behaved when it failed, how much noise the defaults produced untuned, whether the automation actually closed work without you, and how honest the answers were when you asked about leaving.</p>\n    <div class=\"callout\"><p><strong>The one-question version:</strong> ask each vendor to describe, in specifics, the last time their agent failed badly and what they changed because of it. The product whose maker answers plainly is the product whose failures you will hear about while they are still small. That candour predicts your experience better than any feature list ever printed.</p></div>\n    <p>That is how to choose an RMM: not by counting ticks, but by testing the behaviours the ticks conceal, on the estate you actually run.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>An RMM vendor telling you how to buy an RMM should expect scepticism, so we will keep this short and testable. Helios is built around the axes this article says matter: an agent designed to say \"could not check\" rather than pretend, defaults tuned to be quiet until something is wrong, automation through Helio AI that shows its reasoning and its actions, per-device pricing on a rolling term, and your data exportable whenever you want it. We would rather you held every vendor to that standard, us included, because it is the standard the ugly trial reveals anyway.</p>"},{"slug":"slow-computer-troubleshooting","title":"Slow computer troubleshooting: find the real cause before you reimage","meta_description":"Every slow computer has one of a short list of real causes. How to tell disk, memory, background jobs, thermals and network apart before you reimage.","published_at":"2026-08-08T00:00:00","body_html":"<p class=\"lede\">\"It's running slow\" is the most common complaint in IT and the least specific. The machine might be starved of memory, halfway through an antivirus scan, throttling in the afternoon sun, or perfectly healthy and waiting on a struggling server. Slow computer troubleshooting is really a sorting problem: there is a short list of genuine causes, each leaves a distinct signature, and each has a different fix. This guide walks through those signatures in order, whether you support paying clients or the colleague two desks away.</p>\n\n    <p>The reason this matters is that the default response, reimage the machine or just buy a new one, treats every cause identically. Sometimes that works by accident. Often it wastes half a day and the slowness comes straight back, because the cause was never on the machine in the first place. Ten minutes of diagnosis beats four hours of rebuilding almost every time.</p>\n\n    <h2>Start with the symptom, not the machine</h2>\n    <p>Before you touch Task Manager, ask four questions. The answers cut the list of possible causes roughly in half each time:</p>\n    <ul>\n      <li><strong>Slow at what?</strong> \"Everything\" points at disk, memory or thermals. \"Just Outlook\" or \"just the finance system\" points at one application or the service behind it.</li>\n      <li><strong>Slow always, or slow sometimes?</strong> Constant slowness is hardware or configuration. Slowness that comes in bursts is a background job. Slowness at 9am on Mondays is everyone logging in at once.</li>\n      <li><strong>Since when?</strong> \"Since yesterday\" usually means something changed: an update, a new agent, a new browser extension. \"For months, gradually\" points at ageing hardware or accumulating startup load.</li>\n      <li><strong>Just this person?</strong> If three people report it the same week, stop diagnosing the endpoint. The cause is shared: the network, a server, or a change you rolled out.</li>\n    </ul>\n    <p>Write the answers into the ticket. Half the value of a service desk is that the next person does not start from zero, and \"user says slow\" is starting from zero.</p>\n\n    <h2>Slow to start: disk and startup load</h2>\n    <p>If the complaint is about boot time and the first ten minutes of the day, look at the disk first. The single biggest performance divide in any fleet is still spinning hard drive versus SSD, and machines with HDDs are still out there in the long tail of every estate. Check what you are dealing with before anything else:</p>\n    <pre><code>Get-PhysicalDisk | Select-Object FriendlyName, MediaType, HealthStatus</code></pre>\n    <p>If <code>MediaType</code> says HDD on a general-purpose laptop or desktop, you have found your cause and no amount of software tuning will fix it. An SSD swap or a replacement machine is the answer, and everything else in this article is secondary.</p>\n    <p>On SSD machines that are still slow to start, open Task Manager's Startup tab and look at what launches at logon. Estates accumulate startup programs the way lofts accumulate boxes: three chat clients, two cloud sync tools, a printer utility from a printer that left the building in 2023. Each one is small; twenty of them at logon is not. Disable everything that does not need to start automatically, and fix it fleet-wide in policy rather than machine by machine.</p>\n    <p>The other classic signature here is disk sitting at 100% in Task Manager for minutes after logon while the machine catches up on indexing, sync and updates. On an HDD that is crippling; on an SSD it should clear quickly. If it does not, note which process owns the disk column, because that is your next section.</p>\n\n    <div class=\"callout\"><p><strong>Check the real uptime first.</strong> With Windows fast startup enabled, \"shut down\" does not fully restart the machine: only Restart does. Task Manager's Performance tab shows true uptime, and it is routinely weeks. A machine that has not genuinely rebooted in 40 days is slow for reasons no other diagnosis will find. Restart properly, then re-test before you conclude anything.</p></div>\n\n    <h2>Slow all day: memory pressure</h2>\n    <p>If the machine is sluggish constantly, in every app, with fans quiet and disk idle, suspect memory. The signature is easy to read in Task Manager's Performance tab: memory sitting above roughly 85% committed, and the disk ticking over constantly as Windows pages memory out to disk to cope. Paging is the killer. The moment the working set no longer fits in RAM, every application pays a disk-speed penalty for a memory-speed operation.</p>\n    <p>Two things cause it. Either the machine has too little RAM for the year it is living in, 8GB being the current floor for a business Windows machine running a browser, Teams and an office suite together, or one process is eating far more than its share. Sort the Processes tab by memory and look at the top entry. A browser with 60 tabs is the usual suspect; a leaking application that climbs all week and resets on restart is the more interesting one, and worth reporting to whoever owns that app.</p>\n    <p>The fix follows the cause: more RAM if the workload is legitimate, and a conversation or an app fix if it is not. Adding RAM to a machine whose real problem is a memory leak just buys the leak more room.</p>\n\n    <h2>Slow in bursts: background jobs fighting the user</h2>\n    <p>Slowness that arrives in waves, twenty bad minutes and then fine again, is almost never the user's workload. It is something scheduled: an antivirus full scan, a backup job, search indexing, a sync client reconciling thousands of files, or Windows Update downloading and installing in the background. The diagnostic move is correlation. Note the times the user reports, then look at what ran at those times: scan schedules, update history, backup logs.</p>\n    <p>Every one of those jobs is necessary. The failure is scheduling them inside working hours, or letting them all land at once. Move full scans and heavy maintenance to lunchtime or evenings, stagger them across the fleet rather than triggering every machine at 9am, and let updates install at a time you chose rather than whenever the machine feels like it. If updates themselves are the recurring disruption, that is usually a sign they are failing and retrying, which is its own diagnosis: we wrote up <a class=\"inline\" href=\"/blog/why-windows-updates-fail\">how to find the real cause of failing Windows updates</a> separately.</p>\n    <p>One more burst-shaped cause deserves a mention: malware. Cryptominers and other unwanted processes show up exactly like a rogue background job, high CPU from a process nobody recognises, often when the machine is idle. If the process owning the CPU column is not something you installed, treat it as a security question, not a performance one.</p>\n\n    <h2>Slow when hot or unplugged: power and thermals</h2>\n    <p>Some machines are only slow in specific physical conditions, and users rarely volunteer the pattern because they have not spotted it. A laptop that crawls on battery but flies on mains is running a power plan that caps the CPU when unplugged, or is being charged through an underpowered USB-C charger that cannot sustain full performance. A machine that starts the day fine and degrades by afternoon, fans roaring, is thermal throttling: the CPU is deliberately slowing itself to stay within temperature limits.</p>\n    <p>The tell is in Task Manager's CPU graph: a clock speed pinned well below the chip's base frequency while utilisation is high. The causes are unglamorous. Vents blocked by a desk dock or a sofa, fans furred with dust, thermal paste that dried out years ago, or a \"quiet\" fan profile set in firmware. Compressed air and a sensible power plan fix a surprising share of \"this laptop needs replacing\" tickets, at a cost of roughly nothing.</p>\n\n    <h2>Slow, but it is not the computer</h2>\n    <p>A large fraction of slow computer tickets, plausibly the largest, are not about the computer at all. The endpoint is idle: low CPU, plenty of free memory, disk quiet. It is waiting on something at the other end of a wire. The candidates:</p>\n    <ul>\n      <li><strong>The VPN.</strong> If all traffic is backhauled through a VPN before reaching the internet, every cloud app is as slow as that tunnel. The tell: everything is faster off VPN.</li>\n      <li><strong>Wi-Fi.</strong> A weak signal or congested access point makes a fast laptop feel broken. The tell: it is fine on a cable, or fine in a different room.</li>\n      <li><strong>A server or SaaS app.</strong> If the slowness lives in one application for multiple people, diagnose the service, not the laptops of everyone who uses it.</li>\n      <li><strong>Legacy logon plumbing.</strong> Roaming profiles, logon scripts and mapped drives to a wheezing file server produce \"slow computer\" complaints that are really \"slow server\" complaints.</li>\n    </ul>\n    <p>The general test is isolation: same task, same account, different machine or different network. If the symptom follows the network or the service rather than the device, no endpoint fix will help, and a reimage is pure ritual.</p>\n\n    <h2>Slow because it is old: when the diagnosis is the hardware</h2>\n    <p>Sometimes you work through all of the above and the honest answer is that the machine is simply past it. The signature of age is breadth: nothing individually wrong, everything slightly slow, a CPU several generations behind what current software assumes, and a user who has quietly adapted their working day around waiting. That adaptation has a real cost, paid daily, and it is usually far larger than the price difference between replacing a machine at year four versus year six. We covered the economics in our guide to <a class=\"inline\" href=\"/blog/hardware-refresh-cycle\">hardware refresh cycles</a>.</p>\n    <p>This is also the moment to be honest about reimaging. A rebuild fixes exactly one class of cause: accumulated software rot on the endpoint. It does nothing for undersized RAM, a tired disk, blocked vents, the VPN or the file server, which is most of the list above. Reimage when the evidence points at the software layer and you cannot isolate the specific cause, not as a first move. And before you wipe anything, note what the diagnosis showed; a fleet that keeps losing machines to the same cause has a fleet problem, not a machine problem.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Almost every signature in this article is something a monitoring agent can see before anyone opens Task Manager: real uptime, disk type and health, memory pressure over time, the process that owns a CPU spike, patch and scan activity. Helios collects that history from every device it watches, and when a slowness ticket arrives, Helio, the AI layer, runs the investigation you just read about: it checks the device's recent metrics, correlates the complaint with what actually ran, and either fixes the boring causes or hands you a diagnosis instead of a blank ticket. The four questions still matter; you just start with the answers already on screen.</p>"},{"slug":"shadow-it-is-feedback","title":"Shadow IT is feedback: stop blocking it and start reading it","meta_description":"Blocking shadow IT drives it underground and costs you visibility. Why every unsanctioned app is feedback, and how discovery and a faster yes beat prohibition.","published_at":"2026-08-07T00:00:00","body_html":"<p class=\"lede\">Ask almost anyone in this industry what to do about shadow IT and you get the same answer: find it, block it, write a policy forbidding it. We think that answer is wrong, and that it has been making estates less secure for a decade. Whether you look after IT for your own organisation or for clients, every unsanctioned app on the network is a message from your users. Blocking it is choosing not to read it.</p>\n\n    <h2>The standard playbook fails on contact</h2>\n    <p>The orthodox response to shadow IT is prohibition: an acceptable use policy, an application allowlist, a web filter, a stern email. It feels decisive, and on paper the problem goes away.</p>\n    <p>In practice the work the app was doing does not go away, because the deadline that drove someone to it does not go away. The work moves to a personal laptop, a personal Google account, a phone on 4G. You have not reduced shadow IT. You have reduced your <strong>visibility</strong> of it, which is the only lever you actually had.</p>\n    <p>The scale makes prohibition even less plausible. Gartner has put shadow IT at 30 to 40 per cent of IT spending in large organisations, and expects three quarters of employees to acquire, modify or build technology outside IT's oversight by 2027. You are not going to block a majority behaviour. You can only decide whether it happens where you can see it.</p>\n\n    <h2>Every unsanctioned app is a requirements document</h2>\n    <p>Here is the reframe that changes the whole conversation. Nobody adopts an unapproved tool for fun. Somebody risked a telling off, and often spent their own money, because the sanctioned stack failed them in a way they could not get fixed.</p>\n    <p>Marketing living in Canva tells you the design request process is too slow. A project team running itself on Trello tells you the official ticketing tool is too heavy for lightweight work. A free file-sending service in the expense reports tells you your sharing controls block a legitimate everyday need. That is user research, delivered free, with the strongest possible signal of intent: people worked around you to get it.</p>\n    <p>Treat those discoveries as violations and you teach people to hide their needs. Treat them as findings and you get a prioritised list of what to fix in your own stack.</p>\n\n    <h2>The risk was never the app, it is the account</h2>\n    <p>The security argument for blocking sounds robust until you look at what actually goes wrong. Most shadow apps are mainstream SaaS with perfectly serviceable security. The genuine dangers are all properties of <strong>unmanaged use</strong>, not of the software:</p>\n    <ul>\n      <li>No single sign-on or enforced MFA, so one phished password exposes company data.</li>\n      <li>No backup and no retention, so the data is one cancelled free tier from gone.</li>\n      <li>No offboarding. When the person leaves, the account and everything in it leaves with them, a failure mode we covered in our <a class=\"inline\" href=\"/blog/it-offboarding-checklist\">IT offboarding checklist</a>.</li>\n      <li>No inventory entry, so nobody patches it, reviews it or even knows to worry about it.</li>\n    </ul>\n    <p>Every one of those risks is fixed by bringing the app under management, and made permanent by driving it underground. Prohibition does not remove the risk. It guarantees the risk stays in its worst configuration.</p>\n\n    <h2>Discover first, and make amnesty mean it</h2>\n    <p>You cannot triage what you cannot see, so the first real move is discovery, not policy. The signal is already sitting in systems you run: the software inventory your endpoint agent collects, the OAuth consent grants in your Microsoft 365 tenant, the recurring small charges in expense reports.</p>\n    <p>Then declare an amnesty, and mean it. Nobody gets punished for declaring a tool. Each discovery gets one of three honest outcomes: <strong>adopt</strong> it properly, behind SSO and MFA and in the asset register; <strong>replace</strong> it with something that meets the need it revealed; or, rarely, <strong>retire</strong> it, with a written explanation of the specific risk. If most conversations end in a ban, the declarations stop, and you are blind again within a quarter.</p>\n\n    <div class=\"callout\"><p><strong>A useful test:</strong> if your last shadow IT discovery ended with a tool being brought under management, your process is working. If it ended with a warning email and nothing else changing, you did not remove the tool, you removed your last honest source of information about it.</p></div>\n\n    <h2>The root cause is the speed of your yes</h2>\n    <p>The uncomfortable conclusion is that shadow IT is not a user behaviour problem. It is a service level problem, and it is yours. If the sanctioned route to a new tool is a form, a committee and a three week wait, while a credit card takes three minutes, the outcome is designed in. People are not going around IT because they are reckless. They are going around IT because IT is the slow path.</p>\n    <p>So measure your time to yes, and get it under a couple of days for low-risk requests. The teams with the least shadow IT are never the strictest ones. They are the fastest ones, because the sanctioned path wins on convenience, which is the only competition that was ever actually running.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>The discovery step is the part tooling can genuinely shorten. The Helios agent already inventories installed software across every device it monitors, alongside patch state, antivirus and backups, so new and unexpected applications show up in the estate view rather than in an annual audit. With the Microsoft 365 integration connected, sign-in activity in the tenant adds the SaaS side of the picture. What you decide to adopt, replace or retire stays a human call. Helios just makes sure you are deciding from what is actually installed, not from what the policy says should be.</p>"},{"slug":"it-onboarding-checklist","title":"IT onboarding checklist: the steps that get skipped and what they cost","meta_description":"An IT onboarding checklist with the reasoning attached: what to set up before day one, what to prove in the first hour, and what each skipped step costs later.","published_at":"2026-08-06T00:00:00","body_html":"<p class=\"lede\">Most IT onboarding checklists are a list of accounts to create, and creating accounts is the easy half. The expensive half is everything the list quietly assumes: who starts it, how soon, what the new starter should be able to reach unaided, and what happens to the access nobody wrote down. Below is an IT onboarding checklist with the reasoning attached, written for whoever is responsible for the estate, whether that is an in-house team of two, a one person IT department, or an MSP standing up a starter inside a client tenant.</p>\n\n    <h2>Before day one: the week that decides the rest</h2>\n    <p>Roughly a fifth of staff turnover happens in the first month and a half, and a first day spent watching someone else install software is a poor opening impression. Everything in this section should be done before the new starter walks in, which means the clock starts at contract signature, not when someone remembers.</p>\n\n    <h3>1. Trigger the process from the signed contract, not an email</h3>\n    <p>The trigger matters more than the tasks. If onboarding starts with an informal message to whoever is nearest, it starts whenever that person happens to send it, which is usually the Friday before. Hardware has lead times and a device build cannot be compressed to nothing. Without a defined trigger you are not running a process, you are running a favour, and the failure mode is a laptop ordered on the morning someone starts.</p>\n\n    <h3>2. Build the account from a role template, never by copying a colleague</h3>\n    <p>Copying an existing user is the most common shortcut in IT onboarding, and the one that compounds. The colleague being copied carries permissions from three old projects, a temporary elevation nobody removed, and a group whose purpose is long forgotten. Copy them and the new starter inherits all of it on day one. Do that for two years and nobody can explain why anyone has the access they have, which turns every future access review into guesswork.</p>\n\n    <h3>3. Assign licences deliberately and record the decision</h3>\n    <p>Licences get handed out by pattern matching: this person looks a bit like that person, so give them the same. It is how a warehouse supervisor ends up on a premium suite, and how a subscription bill drifts upwards every quarter with no single decision to point at. Record who approved which tier and why, because the reclaim conversation a year later is impossible without it.</p>\n\n    <h3>4. Enrol, encrypt and patch the device before it leaves your hands</h3>\n    <p>A device handed over first and enrolled later is often never enrolled at all. Get it into management, confirm disk encryption is on and the recovery key is escrowed somewhere you can actually reach, check the antivirus is reporting rather than merely installed, and bring it to a current patch level before it ships. Skip this and the newest machine in the estate is the least protected one, invisible to every report you rely on.</p>\n\n    <h2>Day one: the first hour is the whole impression</h2>\n    <p>If the preparation went well, day one is not a build but a verification. Budget half an hour with the person and use it to prove things work rather than to discover they do not.</p>\n\n    <h3>5. Hand over credentials through a channel that cannot be forwarded</h3>\n    <p>A temporary password sent to a personal address, or printed in the welcome pack, lives there permanently and passes through systems you do not control. Use a channel that expires, force a change at first sign-in, and never let the first credential and the second factor travel together. Skipped, this becomes the weakest link in an otherwise well run estate.</p>\n\n    <h3>6. Watch them enrol multi-factor authentication, do not simply require it</h3>\n    <p>The gap between an account going live and its second factor being registered is a real attack window, because whoever enrols first owns the factor. Registering it together, on the spot, closes that window to minutes. Leave it to a prompt the user can defer and a new account with a guessable name sits unprotected for a fortnight. Our <a class=\"inline\" href=\"/blog/microsoft-365-security-checklist\">Microsoft 365 security checklist</a> covers the tenant settings that make this enforceable rather than optional.</p>\n\n    <h3>7. Prove the access instead of assuming it</h3>\n    <p>Walk through it with them: open every application, send a message, join a video call, open the shared folder, reach the line-of-business system. It takes a quarter of an hour. Without it the failures surface one at a time over the next fortnight, each arriving as a ticket that reads like user error and each costing the new starter half a morning.</p>\n\n    <h3>8. Tell them how to reach support and what you will never ask for</h3>\n    <p>New joiners are targeted deliberately: they are eager, they do not yet know who is who, and their arrival is often announced publicly. Give them the real support channel and tell them plainly what IT will never do, which is ask for a password or a code, or message them urgently from an unknown number claiming to be a director. Skip it and their first security decision is made with no information at all.</p>\n\n    <h2>The first fortnight: what the checklist usually forgets</h2>\n    <p>These are the items that get dropped when things are busy, precisely because nothing visibly breaks when they are. The cost lands months later, usually on someone else's desk.</p>\n\n    <h3>9. Record the asset against the person on the day you hand it over</h3>\n    <p>An asset register written from memory at the end of the quarter is fiction. Capture the serial, model, purchase date and holder at the moment of handover, when the information is free. Skip it and you cannot say who has what, kit leaves the building unnoticed for a year, and any <a class=\"inline\" href=\"/blog/hardware-refresh-cycle\">hardware refresh cycle</a> you plan is built on guesses.</p>\n\n    <h3>10. Capture the access that lives outside single sign-on</h3>\n    <p>Directory access you can query later. The rest you cannot: the supplier portal with a shared login, the VPN profile, the shared mailbox, the finance system that has never heard of your identity provider. Write these down as you grant them, because this is the exact list that survives a departure and turns a leaver into a live credential. It is why <a class=\"inline\" href=\"/blog/it-offboarding-checklist\">offboarding fails</a> far more often than onboarding does.</p>\n\n    <h3>11. Settle the local administrator question in writing</h3>\n    <p>Someone will need to install something in week one, and the fastest answer is to grant local admin and move on. Elevation granted informally is permanent elevation, because there is no record of it and therefore no trigger to remove it. Decide the rule, apply it consistently, and if you must elevate someone, put an end date on it in a system that will chase you.</p>\n\n    <h3>12. Book a thirty day review before you close the ticket</h3>\n    <p>Thirty days in, the new starter knows what they still cannot do and which tools they have never opened. That is the one moment when right-sizing access is easy and uncontroversial. Without it, over-provisioning is never corrected and the gaps get worked around with personal accounts, which is how shadow IT actually starts.</p>\n\n    <div class=\"callout\"><p><strong>The honest test:</strong> onboarding is not really judged on day one, it is judged on the day someone leaves. If you cannot produce, in ten minutes, a complete list of what a given person was granted and by whom, the process is creating a security problem rather than preventing one.</p></div>\n\n    <h2>What makes an IT onboarding checklist survive contact with reality</h2>\n    <p>Almost every organisation has a checklist. Far fewer have one that is followed under pressure, and the difference comes down to four things.</p>\n    <p><strong>It lives where the work happens.</strong> A checklist in a document is read once and remembered badly thereafter. It needs to arrive as work, in the same queue as everything else you are accountable for, or it loses every time it competes with an outage.</p>\n    <p><strong>Every item has an owner and a date relative to the start date.</strong> Not \"before they start\" but \"five working days before\". Shared responsibility means nobody is ever late, because nobody was specifically on the hook.</p>\n    <p><strong>It produces a record automatically.</strong> The question an auditor, an insurer or an incoming client asks is not whether you have a process. It is \"show me the last five people you onboarded and what each was given.\" If answering means reconstructing history from mailbox archives, the process exists on paper only.</p>\n    <p><strong>It mirrors your offboarding list.</strong> Anything granted at onboarding and not written down cannot be revoked at offboarding. Building the two as a matched pair, same items in the same order, is the cheapest control available to a small team.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>The practical difficulty with all of the above is that it spans systems: identity in Microsoft 365, the device and its protection state, the asset record, and the ticket tying them together. Helios keeps those in one place, so a starter's device, its patch and antivirus posture, its owner and the onboarding ticket are one object rather than four systems to reconcile, and the protection gaps onboarding tends to create get flagged rather than waited for. None of that removes the discipline above. It means the record writes itself.</p>"},{"slug":"co-managed-it","title":"Co-managed IT: where the model works and where it quietly fails","meta_description":"Co-managed IT is growing fast and most of the advice about it is half right. Where the model genuinely works, where it quietly fails, and what to do instead.","published_at":"2026-08-05T00:00:00","body_html":"<p class=\"lede\">Co-managed IT is the fastest growing shape of the managed services relationship, and the reason is obvious once you say it out loud: it reaches the organisations that already have their own IT people and have no intention of giving them up. The advice about how to run it has settled into a tidy list of rules, and those rules are roughly half right. This piece is for both sides of the arrangement, the provider selling co-managed IT and the internal team living inside it, and it goes through the standard advice one line at a time to separate what holds from what quietly falls over in month four.</p>\n\n    <h2>What the co-managed IT pitch gets right</h2>\n    <p>Start with the part that is true, because it is genuinely true. Kaseya's channel research has 61% of MSP leaders reporting co-managed revenue growing year on year, and that is not a fashion. It is a segment that fully outsourced managed services could never reach, because the buyer was never going to make their own IT staff redundant to sign a contract.</p>\n    <p>The complementary strengths are real too. An internal team of one to five people knows the business: which director never restarts their laptop, which line-of-business application breaks if you touch the print spooler, which invoice run cannot slip. What the same team usually lacks is depth in narrow areas, cover at 2am and during annual leave, tooling that would be absurd to license for forty endpoints, and a second opinion at the moment an incident is going badly. A provider has all four and none of the context. On paper that is a perfect trade.</p>\n    <p>The standard advice that follows is correct in outline too: document who owns what, run both teams on shared tooling, define the escalation path. The trouble is that all three are stated at a level of abstraction where they cannot fail, and the failures all live one level down.</p>\n\n    <h2>Myth one: agree who owns what and the rest follows</h2>\n    <p>The usual implementation is a responsibility matrix drawn up in the sales process. You take patching, servers and backup, we keep user support and the applications. Everyone signs it, everyone feels organised.</p>\n    <p>For <strong>planned work this holds beautifully</strong>. Patch windows, backup verification, starters and leavers, hardware refresh, projects: all of it is scheduled, all of it has an owner before it begins, and a functional split is exactly the right tool. If your co-managed relationship is mostly planned work, the matrix will carry you a long way.</p>\n    <p>It fails on unplanned work, and it fails for a structural reason rather than a discipline one: <strong>incidents do not arrive labelled with their category</strong>. \"Finance cannot get into the ERP\" might be a network problem, an expired certificate, a Microsoft 365 licensing change, a supplier outage or one user's cached credentials. You find out which at minute forty. A matrix that assigns ownership by category is therefore assigning ownership using an answer that nobody has yet. In the gap, both teams either start work or neither does, and the two classic failure modes appear: ticket ping-pong, where each side reclassifies the ticket towards the other, and the double-ticket loop, where the user emails internal IT and also rings the provider, and two technicians now work the same fault from different facts.</p>\n    <p>What actually holds is ownership assigned by <em>state</em> rather than category. Every unclassified incident has one default owner from the second it is raised. Reclassifying it may move the technical work, but it never moves who is accountable for talking to the user, and a handover is an explicit act with the current state written down, not a mention in a chat channel. Deduplication needs to be mechanical as well: same user, same asset, inside the same short window, one ticket.</p>\n\n    <div class=\"callout\"><p><strong>A test worth running this week:</strong> take last month's five worst tickets and ask, of each one, who owned it at minute one and who owned it at resolution. If those two names differ on more than one ticket, your split is by category and it is costing you the first hour of every incident.</p></div>\n\n    <h2>Myth two: co-managed means filling the gaps</h2>\n    <p>\"We flex around your team and fill whatever gaps you have\" is a superb sentence in a sales meeting and a poor operating model. Gaps are, by definition, undefined. They are also infinite, always urgent, and reliably composed of the work nobody in-house wanted to do.</p>\n    <p>Two things go wrong, and they compound. The commercial one is that gap work is interrupt driven and unplannable, which makes it the most expensive kind of work to deliver, and it is almost always sold at a discount on the grounds that \"they do half of it themselves\". The relationship one is worse. Gap-filling puts the provider in permanent implicit competition with the person who recommends whether to renew. If the provider takes credit with end users for fixes the internal team asked for, that team turns hostile inside a quarter, and hostility from that direction is fatal because they control access, information and the renewal conversation.</p>\n    <p>The version that holds is unglamorous: <strong>named services with a defined output</strong>, even where the scope is genuinely broad. \"Patch compliance for all Windows endpoints, monthly evidence pack, exceptions raised within one working day\" survives contact with reality in a way that \"we help with patching\" does not, because it can be delivered, measured and disputed. And the internal IT lead is the client, not the end user. Route the credit to them, present findings to them first, and let them take the report upstairs.</p>\n\n    <h2>Myth three: shared tooling means giving them a login</h2>\n    <p>Everyone agrees both teams should work in the same system. In practice this is usually delivered as a read-only login to the provider's platform, which is not shared tooling at all. It is a window.</p>\n    <p>Read-only has a specific cost: every action the internal team wants becomes a request, every request becomes a ticket, and the ping-pong you removed earlier comes back through a different door. What has to be shared is the state <em>and</em> the permission to act on it. The state that matters is unremarkable and rarely all in one place: device inventory with warranty and age, patch state per device, protection gaps including anything with no working antivirus at all, backup job outcomes, alert history, and a ticket audit trail both teams can read.</p>\n    <p>Two practical points get skipped in the excitement of a new contract. First, do not run two management agents on the same endpoint because the two parties each brought their own. Duplicate patching engines fight over reboots, contradict each other's compliance reporting, and produce the specific misery of a device that is compliant in one console and failing in the other. Second, agree at signature who owns the tenant and the licences, what an export looks like, and how long access persists after notice is served. That conversation is cheap on day one and grim on the day it matters, which is the whole argument of our <a class=\"inline\" href=\"/blog/msp-client-offboarding-checklist\">client offboarding checklist</a>.</p>\n\n    <h2>Myth four: co-managed is the easy version of managed services</h2>\n    <p>It is easier to <em>win</em>. It is harder to run. The overhead of a co-managed relationship is two teams, two calendars, two escalation cultures and a service level commitment measured against a partner who is also doing some of the work, and none of that appears in a proposal that prices co-managed at sixty per cent of the full stack because the client \"does half of it\".</p>\n    <p>Service levels are where this bites first. If your first response depends on the internal team granting access, providing a decision or confirming a change window, then either that waiting time is excluded from the clock or you should not be promising the number. Say which, in writing, before the first breach rather than after it. Our guide to <a class=\"inline\" href=\"/blog/msp-sla-response-times\">setting SLA response targets you can actually hit</a> applies here with one extra rule: a co-managed SLA needs a stated dependency clause or it is fiction.</p>\n    <p>Pricing follows the same logic. The cost driver in a co-managed agreement is the surface you are responsible for, not the number of tickets you happen to receive, because coverage costs the same whether or not anyone rings. Price the estate you monitor and the named services you agreed, add blocks for project work, and treat genuine out-of-hours cover as its own line rather than a goodwill gesture. \"We sleep, you do not\" is a real service and it is worth what it costs.</p>\n    <p>Ninety days in, the model is working if the internal team's own backlog is shrinking rather than the provider's, if escalations arrive with context attached instead of a forwarded email, if no incident this month was worked twice, and if the internal lead is the person presenting results to the board. If instead the provider is quietly becoming the second line for everything, you have not built a partnership, you have built an <a class=\"inline\" href=\"/blog/one-person-it-department\">understaffed internal team</a> with an invoice attached.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios is multi-tenant by design, so a provider and an internal team can work the same estate without either one being reduced to a spectator: the same device inventory, patch state, protection gaps, backup outcomes and ticket history, with roles that let the internal team act rather than only look. Helio, our AI layer, triages incoming tickets and runs device investigations before either team picks the ticket up, which is the part of the first hour that both sides currently spend arguing about ownership. One agent reports patching, antivirus and backup state, so there is one version of the truth to disagree with.</p>"},{"slug":"do-phishing-simulations-work","title":"Do phishing simulations work? The evidence says barely","meta_description":"A trial of 19,500 employees found phishing training barely moves click rates. Why phishing simulations keep failing, and what stops phishing instead.","published_at":"2026-08-03T00:00:00","body_html":"<p class=\"lede\">Do phishing simulations work? Almost everyone in security behaves as if the answer is obviously yes. Compliance frameworks expect them, cyber insurers ask about them, and a large awareness-training industry depends on them. Yet the best evidence we now have says the honest answer is barely, and whoever you run IT for, an MSP's clients or your own colleagues, that should change where your hours and budget go.</p>\n\n    <h2>The result the industry does not want to discuss</h2>\n    <p>In 2025, researchers at UC San Diego Health published one of the largest real-world tests of phishing training ever run: a randomised trial across roughly 19,500 employees, with ten simulated campaigns over eight months. The headline finding was blunt. There was no significant relationship between having recently completed the organisation's annual security training and the likelihood of falling for a phishing email. The embedded lessons that appear after someone clicks a test lure moved failure rates by about 1.7 percentage points. That is the entire measured benefit of the thing your whole programme is built around.</p>\n    <p>The behavioural detail explains why. More than three quarters of staff spent under a minute on the training material, and between a third and a half closed the page almost immediately. People are not learning from this content because they are not reading it, and no amount of gamification has changed that.</p>\n    <p>This was not a one-off. A 2021 ETH Zurich study followed around 14,000 employees of a large company for fifteen months and found that embedded training after a failed test did not reduce future clicks, with some evidence it made users more likely to fall for later attacks, possibly by leaving them feeling protected. Two large, careful, independent studies, years apart, pointing the same way.</p>\n\n    <h2>Someone always clicks, and the maths does the rest</h2>\n    <p>Even if training worked better than it does, simulations would still be aimed at the wrong number. A phishing programme tries to lower your average click rate. An attacker does not care about your average. They need one person, on one bad day, reading one well-timed email on a phone between meetings.</p>\n    <p>Suppose training gets your click rate down to a widely envied three percent. Across 200 users, one campaign still lands six footholds. Run a campaign a month and the question is not whether someone clicks but how quickly. The UC San Diego data shows exactly this compounding: in the first month about one in ten employees clicked a lure, but by month eight more than half had clicked at least once. Individually rare events, repeated across enough people and enough emails, become a certainty.</p>\n    <p>You cannot train your way out of that arithmetic. Every user you support will eventually click something. A security posture that depends on that never happening is not a posture, it is a hope.</p>\n\n    <h2>A click rate measures your lure, not your people</h2>\n    <p>Here is the quieter problem with phishing simulations: the metric they produce is adjustable at will. Send an obviously clumsy lure and your click rate looks wonderful. Send a convincing one about a bonus scheme or a missed parcel and it looks terrible. The number the board sees is mostly a property of the test, not of the workforce.</p>\n    <p>That makes the click rate almost uniquely bad as a security KPI. It can be steered up to justify budget or down to show progress, and neither movement tells you anything about whether a real, targeted attack would succeed. A metric you can set to whatever you need it to be is not a measurement. It is theatre with a percentage sign.</p>\n    <p>The same distortion runs through the industry's favourite statistics. Vendor case studies showing dramatic improvement almost always compare an easy later campaign with a hard early one, or measure users who were about to improve anyway. The controlled studies, the ones with randomisation and a control group, are the ones that keep finding almost nothing. When the rigour goes up, the effect goes away. That pattern has a name in every other field, and it is not a flattering one.</p>\n\n    <h2>The costs are real even if the benefit is not</h2>\n    <p>None of this would matter much if simulations were free. They are not. There is the direct cost of the platform and the thousands of staff-hours spent on lessons the data says are skimmed in under a minute. There are the trust costs: employees learn that IT sends them traps, legitimate emails get reported and sit in quarantine while someone in finance chases an unpaid invoice, and the fake-bonus-email genre has produced public HR disasters that outlived any security benefit.</p>\n    <p>The biggest cost is the least visible. A phishing programme lets an organisation feel it has dealt with phishing. That feeling absorbs budget, attention and board goodwill that should have gone to controls which actually determine whether a click matters. Training is not just weakly effective. It is a sedative.</p>\n\n    <h2>Assume the click: the controls that actually stop phishing</h2>\n    <p>The alternative is not resignation, it is engineering. Design the estate so that a click is survivable, and the tail-risk arithmetic flips to your side:</p>\n    <ul>\n      <li><strong>Phishing-resistant MFA.</strong> Passkeys and FIDO2 keys cannot be tricked into signing in to a fake site, and they do not push approval prompts that tired thumbs accept. This one control removes the value of most credential lures outright.</li>\n      <li><strong>Close the tenant's side doors.</strong> Legacy authentication, unrestricted OAuth app consent and silent forwarding rules are how a phished password becomes a persistent breach. Our <a class=\"inline\" href=\"/blog/microsoft-365-security-checklist\">Microsoft 365 security checklist</a> covers the ten settings to fix first.</li>\n      <li><strong>Patch the software the click lands in.</strong> A malicious page or attachment needs a vulnerable browser or reader to become code execution. Keeping those current is unglamorous and hugely effective, and as we have written before, <a class=\"inline\" href=\"/blog/third-party-patching-for-msps\">third-party patching is harder than Windows Update</a> and needs its own mechanism.</li>\n      <li><strong>Least privilege.</strong> A payload that runs without local admin rights lands in a puddle, not a lake.</li>\n      <li><strong>Fast reporting and response.</strong> A one-click report button, and a desk that can pull a lure from every mailbox in minutes. Time-to-quarantine is a real security metric. Click rate is not.</li>\n    </ul>\n    <p>Every item on that list is boring, measurable and indifferent to human error. None of them can be skimmed in under a minute, forgotten by Friday or quietly closed like a training tab, because none of them lives in anyone's head. That is precisely why each is worth more than another round of training videos, and why the teams with the calmest phishing story are usually the ones that spent their awareness budget here instead.</p>\n\n    <div class=\"callout\"><p><strong>A fair test:</strong> if your organisation spends more per year on awareness training than on making MFA phishing-resistant and keeping browsers patched, your phishing budget is upside down. Fix the controls first; they work on the day someone clicks, which training demonstrably does not.</p></div>\n\n    <h2>What simulations are still good for</h2>\n    <p>This is an argument for demotion, not abolition. Keep a light simulation programme, but change what it measures. Report rate and time-to-report are worth tracking, because a workforce that forwards suspicious emails quickly genuinely shortens incidents. An occasional exercise that tests your response pipeline, how fast the lure was reported, triaged and purged from mailboxes, tells you something real about your defences. And never publish league tables or punish clickers: the moment reporting feels risky, your best early-warning system goes silent.</p>\n    <p>So, do phishing simulations work as a way to stop people clicking? Barely, and the largest studies we have say so plainly. Treat them as a smoke detector test for your response process, spend the reclaimed money on controls that assume the click, and your next real phish will meet an estate that does not need anyone to be perfect.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios is built for the assume-the-click side of this argument. The platform continuously checks the controls that decide whether a phish matters: protection-gap detection flags devices where Defender or another AV is missing or unhealthy, third-party patching keeps browsers and readers current alongside Windows, and the Microsoft 365 integration surfaces risky tenant settings before an attacker finds them. When users do report something suspicious, Helio's auto-triage helps the desk separate the real lure from the noise quickly, which is the metric that actually counts.</p>"},{"slug":"msp-pricing-models","title":"MSP pricing models: per-device, per-user or flat fee?","meta_description":"An honest comparison of the three MSP pricing models, per-device, per-user and flat fee: how each one fails, and the hybrid most MSPs should actually run.","published_at":"2026-08-02T00:00:00","body_html":"<p class=\"lede\">Every conversation about MSP pricing models eventually collapses into the same three options: charge per device, charge per user, or charge one flat fee for the whole estate. All three are defensible, all three are in daily use, and all three fail in specific, predictable ways. This is a comparison of how each one actually behaves over a contract's lifetime, and, because a comparison that ends in \"it depends\" is useless, an actual recommendation at the end.</p>\n\n    <h2>What a pricing model is really for</h2>\n    <p>Before comparing the options, it is worth being precise about the job. A pricing model is not just a way of calculating an invoice. It does three things at once:</p>\n    <ul>\n      <li><strong>It allocates risk.</strong> Somebody absorbs the difference between the fee and the real cost of service in any given month. The model decides whether that is you or the client.</li>\n      <li><strong>It shapes client behaviour.</strong> Whatever you charge for, clients will try to have less of. Charge per ticket and they stop reporting problems. Charge per device and old machines never get retired from the count, just hidden.</li>\n      <li><strong>It sets the tone of every future conversation.</strong> A model the client cannot predict produces a billing argument every month. A model they can predict produces a renewal.</li>\n    </ul>\n    <p>The reason MSPs argue about pricing models so much is that no model does all three jobs well. Each one trades precision for simplicity somewhere, and the right question is not \"which model is correct\" but \"which trade-offs can my operation and my clients actually live with\".</p>\n\n    <h2>Per-device pricing: precise, and precisely wrong</h2>\n    <p>Per-device is the oldest model, and its logic is honest: your tooling costs scale per endpoint, your patching and monitoring effort scales per endpoint, so your price should too. A workstation costs one rate, a server a higher one, network kit something in between. The invoice is a stock count.</p>\n    <p>Its strengths are real. It maps almost perfectly onto your cost base, so margin per client is easy to see. It is trivially auditable: the client can count their own machines. And for estates with unusual shapes, a warehouse with forty shared terminals and nine staff, a clinic with more diagnostic machines than people, it is the only model that prices the actual work.</p>\n    <p>But per-device taxes exactly the wrong behaviour. The industry has spent a decade telling businesses that security means every laptop enrolled, every mobile managed, every server monitored. Under per-device pricing, each of those good decisions raises the client's bill, so clients quietly resist them. The unmanaged personal laptop that never made it onto the invoice is also the one that never got patched, and it is still your incident when it goes wrong. A model that gives your clients a financial reason to hide endpoints from you is working against your own security posture.</p>\n    <p>It also ages badly. Modern users average two to three devices each, and that number only rises. A model that meters the thing that is multiplying makes every renewal conversation a price rise conversation, which is partly why industry surveys now put per-device somewhere under one MSP in five, and falling.</p>\n\n    <h2>Per-user pricing: the default, for good reasons and one bad one</h2>\n    <p>Per-user charges a single monthly rate for each person supported, covering whatever devices that person reasonably uses. It has become the industry default, and mostly on merit:</p>\n    <ul>\n      <li><strong>It matches how clients think.</strong> A managing director knows their headcount to the person. They could not tell you their device count within twenty percent. An invoice denominated in people is one they can sanity-check in their head, which is worth more goodwill than any discount.</li>\n      <li><strong>It rewards the right behaviour.</strong> Enrolling a second or third device per user costs the client nothing extra, so nothing pushes back against full coverage. Your security posture and your pricing model finally point the same way.</li>\n      <li><strong>It tracks the licensing world.</strong> Microsoft 365, most security tooling and most SaaS is already per-user. Your cost of goods and your revenue rise and fall together with headcount.</li>\n    </ul>\n    <p>The bad reason it became the default is that it is easy to copy without doing the sums underneath. Per-user is an averaging model: it works when the average user's device load and support demand match what you priced. Sign a client whose \"users\" each carry a laptop, a desktop, two mobiles and a tablet, or a client with heavy shared infrastructure and few people, and the average breaks in your direction. Per-user also hides servers entirely. A server is not a user, and folding a demanding hypervisor estate into a per-head rate is how MSPs end up doing their hardest work for free.</p>\n    <div class=\"callout\"><p><strong>The test for a per-user rate:</strong> take your three most expensive clients, divide real monthly effort and tooling cost by their headcount, and compare it to what you charge. If you have never done this arithmetic, your per-user rate is a guess wearing a spreadsheet's clothes.</p></div>\n\n    <h2>Flat-fee pricing: the best sales pitch and the worst discipline</h2>\n    <p>The third model drops the metering entirely: one fixed monthly fee for the whole environment, everything included. It is the easiest model to sell, because it is the one clients actually want. No counting, no surprises, one line on the budget. It is also the purest expression of what managed services are supposed to be: you are buying an outcome, a working estate, not a basket of units.</p>\n    <p>For the MSP, flat-fee done well is the most profitable model on this list, because it decouples revenue from effort. Every hour of toil you automate away widens the margin instead of shrinking the invoice. It gives you the strongest possible incentive to prevent problems rather than bill for them.</p>\n    <p>Done badly, it is the fastest way to go broke slowly. A flat fee is an underwriting exercise: you are insuring the client against their own IT demand, and if you set the premium without data, the client with the flaky server room and the growth spurt will consume your margin invisibly, month after month, with no mechanism in the contract to correct it. Flat-fee also invites silent scope creep. When nothing is metered, nothing is out of scope by default, and two years in you are supporting an acquired subsidiary, a second office and a fleet of tablets that were never priced at all.</p>\n    <p>Flat-fee is not really a different model so much as a bet that you know the client's demand curve better than they do. Sometimes that bet is right. It should never be the opening offer to a client whose estate you have not yet measured.</p>\n\n    <h2>The comparison that matters: how each model fails</h2>\n    <p>Feature lists make the three models look interchangeable. Their failure modes do not:</p>\n    <ul>\n      <li><strong>Per-device fails at the edges of the count.</strong> Its characteristic failure is the unmanaged endpoint: the machine kept off the invoice that becomes your breach anyway. Secondary failure: renewal friction, because the metered unit is the one multiplying.</li>\n      <li><strong>Per-user fails in the averages.</strong> Its characteristic failure is the client whose per-head demand quietly exceeds the rate, invisible until you measure effort per client. Secondary failure: unpriced servers and infrastructure.</li>\n      <li><strong>Flat-fee fails in the drift.</strong> Its characteristic failure is scope creep with no meter to catch it, discovered only when a profitable contract has become an unprofitable one with the same signature on it.</li>\n    </ul>\n    <p>Notice what the three failures have in common: none of them show up on the invoice. Every pricing model fails silently, which is why the model matters less than the measurement underneath it. An MSP that tracks devices, users and effort per client can run any of these models and correct course at review time. An MSP that tracks none of them is guessing under all three, just with different notation. If you price per user, you still need the device count; if you price flat, you need both, plus effort. This is also why pricing belongs on the agenda of your <a class=\"inline\" href=\"/blog/msp-qbr-guide\">quarterly business reviews</a>: the QBR is where the drift gets caught while it is still a conversation rather than a dispute.</p>\n\n    <h2>Our honest recommendation</h2>\n    <p>If you want a straight answer: <strong>run per-user as your base, price servers and infrastructure per device on top, and treat flat-fee as something a client graduates to, not something they start on.</strong></p>\n    <p>Per-user as the base because its incentives are the least destructive. The failure modes of per-device pricing damage the client's security and your relationship; the failure modes of per-user pricing damage only your margin, and margin failures are fixable with measurement and a rate review. Add a per-device line for servers, hypervisors and network infrastructure because those genuinely are units of work with no user attached, and folding them into a per-head rate is the single most common way good per-user pricing goes bad.</p>\n    <p>Reserve pure flat-fee for clients where you have at least a year of real data on devices, tickets and effort, and where the relationship justifies underwriting their demand. At that point it is a genuinely better model for both sides. Before that point it is a guess, and the discovery work in a proper <a class=\"inline\" href=\"/blog/msp-client-onboarding-checklist\">client onboarding</a> is exactly what turns the guess into a quote.</p>\n    <p>Two caveats, honestly held. If your client base is dominated by device-heavy, people-light estates, per-device is not legacy thinking, it is the correct model for you; run it and ignore the surveys. And whichever model you choose, write the assumptions into the contract: devices per user, server counts, what triggers a re-price. The model is the headline; the assumptions are the contract.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Every recommendation above leans on the same prerequisite: knowing, per client, how many users and devices you actually support and where the effort goes. Helios keeps that inventory live across every client tenant, devices, users, servers and their state, so a rate review starts from real counts rather than a walk-around, and the client whose demand has outgrown their rate shows up in the data before they show up in your margin. And because Helio's auto-triage and auto-healing take repeat toil off the desk, the gap between a fixed fee and your real cost to serve moves in your favour, which is the whole economic point of fixed pricing done well.</p>"},{"slug":"how-fast-should-you-patch","title":"How fast should you patch? Faster than your rings allow","meta_description":"Exploits land in days; most estates patch in months. Why staged deployment rings are risk theatre for smaller fleets, and the faster patching cadence to run instead.","published_at":"2026-08-01T00:00:00","body_html":"<p class=\"lede\">Ask how fast you should patch and the industry hands you a ritual: a pilot ring, a 72-hour wait, a broader ring, another wait, then the rest of the estate, three or four weeks after release. That advice was written for fleets of fifty thousand machines and a team that genuinely watches the pilot ring. For almost everyone else, we think it is the riskier choice, and this piece is the case for patching much faster.</p>\n\n    <h2>The maths the standard advice ignores</h2>\n    <p>Deployment rings exist to manage one risk: a bad patch breaking things. They manage it by adding time. What the ritual never prices in is what that time now costs, because the other side of the ledger has changed completely.</p>\n    <p>The median time from a vulnerability being disclosed to it being exploited in the wild is now measured in days, often under five, and a growing share of vulnerabilities are exploited before a patch even exists. Meanwhile the average organisation takes more than 60 days to remediate a critical vulnerability, and roughly six in ten breaches involve a vulnerability for which a patch was already available. In several recent incidents, exploitation began within 24 hours of a vulnerability being added to CISA's Known Exploited Vulnerabilities catalogue.</p>\n    <p>Put those numbers side by side and the conclusion is uncomfortable. The window between a patch being released and your estate installing it is not slack in the schedule. It is the attack window, and it is where most real-world compromises now happen. Every 72-hour pause between rings is not caution. It is three more days spent inside the exploited zone, multiplied by every machine still waiting.</p>\n\n    <h2>A ring nobody watches is just a delay</h2>\n    <p>Here is the question that decides whether your rings are a control or a costume: when your pilot ring finished last month, what did anyone actually check?</p>\n    <p>At enterprise scale there is a real answer. Application owners run smoke tests, a helpdesk watches ticket volumes from the pilot cohort, and advancement to the next ring is a decision someone makes against evidence. That is testing, and it earns its delay.</p>\n    <p>In most estates of 30, 100 or even 500 machines, the honest answer is nothing. Nobody ran a test plan. Nobody compared failure rates. The pilot ring \"passed\" because nobody phoned to complain, which is the same evidence you would have had by deploying everywhere. You have adopted the enterprise pattern's delay without its validation, which means you are carrying all of the cost and none of the benefit. A ring nobody watches is not a safeguard. It is a queue.</p>\n    <p>The ceremony compounds quietly, too. A pilot ring on Tuesday, 72 hours of not looking at it, a broader ring, 72 more, a change freeze on Fridays and then the weekend, and a fix released on the 10th reaches the last laptop on the 4th of the following month. Nobody decided to accept three and a half weeks of exposure. The process decided it, one reasonable-sounding pause at a time.</p>\n    <p>The strongest recent argument for rings does not even survive contact with its own example. The most catastrophic bad update in years, the CrowdStrike outage of July 2024, was not a Windows patch. It was a security vendor's content update, pushed through a channel that ignored customers' staging policies entirely. Organisations that had carefully configured n-1 and n-2 update rings were flattened alongside everyone else. The disaster everyone cites to justify slow patching is one that rings demonstrably failed to prevent, while the breaches that known vulnerabilities cause every week rarely make the news at all.</p>\n\n    <h2>Weigh a bad patch against a breach, honestly</h2>\n    <p>None of this claims bad patches do not happen. They do, a few times a year, and occasionally a nasty one. The argument is that the two risks are wildly asymmetric, and the ritual treats them as equal.</p>\n    <p>A bad patch is loud, fast and reversible. Machines blue-screen, a printer driver dies, an application refuses to start, and you know within hours because users tell you. Windows can uninstall a cumulative update, Microsoft ships Known Issue Rollback fixes for widespread problems, and you can pause a deployment mid-flight. If your <a class=\"inline\" href=\"/blog/backup-monitoring-for-msps\">backups are monitored and restorable</a>, even the worst case is a bad day, not a bad quarter.</p>\n    <p>A breach through a known, unpatched vulnerability is the opposite in every respect. It is silent, it dwells for weeks, and nothing about it can be rolled back: not the exfiltrated data, not the ransomware deployment, not the incident-response invoice. You cannot uninstall a compromise.</p>\n    <div class=\"callout\"><p><strong>The honest comparison:</strong> patching fast risks a rare, visible, recoverable failure. Patching slow risks a common, invisible, unrecoverable one. Choosing the second because the first feels more embarrassing is optimising for how the incident will look, not for how much it will cost.</p></div>\n\n    <h2>How fast should you patch, then?</h2>\n    <p>Fast enough that the release-to-deployed gap is measured in days. A cadence we would defend for a typical estate looks like this:</p>\n    <ul>\n      <li><strong>Security updates: auto-approved, deployed within 72 hours of release.</strong> Not after a two-week soak. If a vulnerability is in the Known Exploited Vulnerabilities catalogue, deploy immediately, everywhere, including servers.</li>\n      <li><strong>Keep a canary, but size it in hours, not weeks.</strong> A handful of machines that install on day zero, advanced automatically the next day unless failures actually appear. The point of a canary is automated observation, not a fortnight of ceremony.</li>\n      <li><strong>Servers get care, not delay.</strong> A snapshot, a maintenance window, a tested restore path, and then the patch, days after release rather than months. Care is preparation; delay is just exposure with paperwork.</li>\n      <li><strong>Third-party applications run continuously.</strong> Browsers and readers are the most exploited software you manage and their <a class=\"inline\" href=\"/blog/third-party-patching-for-msps\">patching sits outside Windows Update entirely</a>, so it should not wait for a monthly cycle at all.</li>\n    </ul>\n    <p>One distinction keeps this from being reckless: the argument is about security updates, not feature upgrades. Defer a Windows feature update for months if you like, since nothing exploits your not having the newest Start menu, and the compatibility risk there is real. The mistake is letting the caution appropriate for feature changes set the tempo for security fixes, because the two ride entirely different risk curves.</p>\n    <p>The prerequisites are unglamorous: per-device visibility of what failed to install, reboot discipline so patches actually take effect, and a restore path you have tested. If a machine reports a failure, <a class=\"inline\" href=\"/blog/why-windows-updates-fail\">find the real cause</a> instead of widening the delay for everyone else. And measure the thing that matters: not whether you have rings, but how many days pass between a patch being released and 95% of your estate having it. Whether you look after your own organisation's machines or a dozen clients' estates, that one number is your actual exposure, and for most teams reading this, shrinking it will do more for security than any other change this year.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>The cadence above only works if the boring parts are automatic, which is why Helios treats them as platform work rather than technician work. Patch policies auto-approve and deploy on the schedule you set, Windows and third-party alike, and every device reports per-update success or failure rather than a fleet-wide green tick. Because the same agent watches health continuously, a canary group is genuinely observed: if failures spike after a rollout, you see it in hours and can pause, and Helio's auto-healing picks up the routine failures without a ticket. Speed is safe when the feedback loop is real.</p>"},{"slug":"it-offboarding-checklist","title":"IT offboarding checklist: how a leaver's account becomes a breach","meta_description":"Around one in four ex-employees can still log in somewhere. Where IT offboarding really breaks, traced to root cause, and the checklist that closes every door.","published_at":"2026-07-31T00:00:00","body_html":"<p class=\"lede\">Every IT offboarding checklist exists because of the same incident, repeated endlessly: months after someone leaves, one of their accounts logs in. Surveys keep finding that around one in four ex-employees can still access something they should not, and a meaningful share of breaches trace straight back to a departure that was never fully closed out. This is a post-mortem of that incident, whether you run IT in-house or look after other businesses' estates, and the checklist that stops it happening to you.</p>\n\n    <h2>Six months after the leaving card, a login</h2>\n    <p>The incident almost always looks the same. Someone left in the spring: a sales manager, a contractor, a systems administrator. There was a leaving card, a handover document that was mostly links, and a ticket that said \"please disable Jane's account\", which somebody did. In the autumn, something odd surfaces. A customer list turns up at a competitor. A mailbox rule is quietly forwarding invoices. A VPN session appears from an IP address nobody recognises. The sign-in logs are checked, and there it is: a live login from an account belonging to a person who has not worked there for six months.</p>\n    <p>Sometimes the leaver themselves is behind it, and disgruntled-insider cases make the headlines because they end in court. More often it is worse in a mundane way: the credentials were reused somewhere that got breached, and an attacker simply tried them everywhere. The account they landed in had no active human watching it, no manager noticing odd behaviour, and often no MFA, because the phone that held the authenticator was handed back months ago while the account stayed live.</p>\n    <p>The immediate reaction is always to ask who forgot. That is the wrong question, and it is why the incident recurs. Nobody forgot. Everybody did exactly what the process asked of them. The process asked for the wrong things.</p>\n\n    <h2>Tracing it back: the trigger lived in the wrong department</h2>\n    <p>Run the post-mortem honestly and the first root cause is almost never technical. It is organisational: <strong>leaving a company is an HR event, but access is an IT asset</strong>, and in most businesses there is no reliable wire connecting the two. HR processes the departure for payroll and paperwork on their timeline. IT finds out by email, by corridor, or occasionally not at all. If you are an MSP, add another hop: your client's HR tells your client's office manager, who may or may not tell you, days later.</p>\n    <p>So the first failure is the trigger. The ticket that says \"disable Jane's account\" arrives late, or never, and its timing depends on someone remembering to send it. For resignations with notice that is sloppy; for terminations it is dangerous, because the highest-risk departures are exactly the ones where access must end at the moment the person is told, not at the end of a payroll cycle.</p>\n    <p>The second failure is the wording. \"Disable the account\", singular, reflects a mental model in which a person has one account. That was true in 2005. Today the directory account is one door among many, and disabling it closes the biggest door while leaving the side doors not just unlocked but unwatched. The ticket got closed because the thing the ticket named got done. The gap between \"the ticket is closed\" and \"the person has no access\" is where the breach lives.</p>\n\n    <h2>One leaver, a dozen doors</h2>\n    <p>Inventory what a typical employee actually accumulates and the singular \"account\" becomes obviously inadequate. A leaver commonly walks away still holding some of these:</p>\n    <ul>\n      <li><strong>SaaS accounts outside single sign-on.</strong> The design tool, the expenses app, the CRM seat someone bought on a card. If it was set up with an email address and a password rather than through your identity provider, disabling the directory account does not touch it.</li>\n      <li><strong>Active sessions and tokens.</strong> Disabling an account stops new logins. Sessions already open on a home laptop, and refresh tokens on a personal phone, can keep working for hours or days unless you revoke them explicitly.</li>\n      <li><strong>Mailbox rules and forwarding.</strong> A forwarding rule set up years ago, legitimately or not, keeps forwarding after the person leaves. Our <a class=\"inline\" href=\"/blog/microsoft-365-security-checklist\">Microsoft 365 security checklist</a> covers why forwarding deserves standing scrutiny; offboarding is the moment it bites.</li>\n      <li><strong>Shared credentials.</strong> The router password, the social media logins, the admin account \"everyone\" uses. You cannot disable knowledge. Anything the leaver knew has to be rotated, and first you have to know what they knew.</li>\n      <li><strong>Remote access.</strong> VPN certificates, remote desktop gateways, remote support tools installed during some old troubleshooting session. These are the doors attackers like best, because nobody watches them.</li>\n      <li><strong>Devices and data.</strong> The laptop in a cupboard at their house, the phone with company email on it, the personal OneDrive the handover files got synced to in the final week.</li>\n      <li><strong>Keys they created.</strong> Especially for technical leavers: API keys, service account passwords, automation credentials made under their own identity years ago and still in production.</li>\n    </ul>\n    <p>This is the third root cause: <strong>the only complete map of a person's access existed in that person's head</strong>, and it left when they did. No checklist fixes that on the last day. The fix is keeping the map continuously, which is an inventory problem, not an offboarding problem.</p>\n\n    <h2>The IT offboarding checklist, in three passes</h2>\n    <p>A checklist that actually prevents the incident is structured by time, because the actions have different deadlines. One pass on the last day is not enough; three passes are.</p>\n    <h3>Before the last day</h3>\n    <ul>\n      <li>Open the offboarding ticket the moment the departure is confirmed, not when it takes effect.</li>\n      <li>Build the access map: assets assigned, group memberships, admin roles, SaaS seats, shared credentials known, API keys and service accounts created.</li>\n      <li>Agree data handover with the manager: which files, which mail, who inherits what.</li>\n    </ul>\n    <h3>Day zero: the last day, or the moment of termination</h3>\n    <ul>\n      <li>Disable the identity in Microsoft Entra ID or Active Directory, and revoke all active sessions and refresh tokens, not just future logins.</li>\n      <li>Remove MFA methods and app passwords registered to the account.</li>\n      <li>Disable VPN and remote access, and uninstall or disable remote support tools on their devices.</li>\n      <li>Rotate every shared credential on the access map.</li>\n      <li>Recover the hardware, or remote wipe what does not come back the same day.</li>\n    </ul>\n    <h3>The first week and day 30</h3>\n    <ul>\n      <li>Convert the mailbox to a shared mailbox or delegate it. Review it for forwarding rules before anyone relies on it.</li>\n      <li>Transfer file ownership, reclaim licences, remove the account from distribution lists and Teams.</li>\n      <li>Close SaaS accounts outside SSO, and rotate any API keys and service credentials the leaver had reach into.</li>\n      <li>At day 30, verify: check sign-in logs for any successful authentication since day zero, sweep SaaS billing and the password manager for accounts you missed, then archive or delete per your retention policy.</li>\n    </ul>\n    <div class=\"callout\"><p><strong>Disable, never delete, on day zero.</strong> Deleting an account destroys the evidence trail, orphans files and mail, and can break anything running under that identity. Disable immediately, revoke everything, and delete only after the day-30 verification pass, when retention allows.</p></div>\n\n    <h2>The change that prevents it: a trigger, an owner, and proof</h2>\n    <p>The checklist above is necessary but not sufficient, because the original failure was not a missing list. It was a missing trigger, a singular mental model, and a map that lived in someone's head. The preventive change has three parts, and none of them is glamorous.</p>\n    <p><strong>First, wire the trigger.</strong> Departure confirmed in HR must create the IT offboarding ticket automatically, or by an agreed same-day handoff that someone is accountable for. If you are an MSP, put this in the contract: the client notifies you of leavers within one working day, and immediately for terminations. Then measure it. Time from departure to identity disabled is the one metric that predicts whether this incident can happen to you, and for privileged accounts the acceptable answer is minutes.</p>\n    <p><strong>Second, give every departure a named owner</strong> who closes the ticket only when every task is done, not when the biggest one is. The checklist lives as tasks in the ticket, each with evidence attached: the disabled state, the wipe confirmation, the rotation record. \"Done\" becomes checkable by someone else, which is the entire point of a checklist.</p>\n    <p><strong>Third, verify with data rather than memory.</strong> The day-30 log review catches what the last day missed. A quarterly sweep comparing your directory and your asset inventory against the actual list of current staff catches what day 30 missed. Stale accounts are findable by machine: no sign-in for 45 days, licence still assigned, owner not in payroll. If your tooling keeps a live inventory of users, devices and installed software, that sweep is a saved report, not an afternoon. And the same discipline applies at bigger scale when an entire client relationship ends; our <a class=\"inline\" href=\"/blog/msp-client-offboarding-checklist\">client offboarding checklist</a> is the estate-sized version of this article.</p>\n    <p>Do these three things and the opening incident becomes structurally difficult. The account is disabled the day the person leaves because the trigger fired. The side doors are closed because the map existed before the last day. And anything that slipped through is caught within a month by a review that looks at logs, not at recollections.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios does not do your offboarding for you, but it removes the two ingredients the incident depends on: the missing map and the missing verification. The asset inventory ties devices, installed software and users together, so \"what did they have\" is a lookup rather than an archaeology project. The Microsoft 365 integration surfaces account and licence state alongside the endpoint view, ticket checklists give each departure an owner and an evidence trail, and Helio can flag the pattern that should never occur: activity attributed to a user the estate no longer contains.</p>"},{"slug":"why-windows-updates-fail","title":"Why Windows updates fail: how to find the real cause","meta_description":"The same machines fail Windows updates every month for a reason. How to tell disk space, reboot loops, broken components and dead WSUS apart, and fix each.","published_at":"2026-07-30T00:00:00","body_html":"<p class=\"lede\">Patch reports rarely show a disaster. They show a plateau: 92% compliant this month, 91% last month, and the same few devices in the failed column every time. Working out why Windows updates fail is mostly about reading that failed column properly, because a failed update is a symptom with about five common causes, each with its own tell and its own fix. That is true whether the machines belong to clients or to your own colleagues.</p>\n\n    <h2>The symptom: a compliance number that will not move</h2>\n    <p>Windows Update failures are not evenly spread. On a healthy estate, most machines patch themselves without ever being noticed, and the failures concentrate on repeat offenders: a machine that fails an update this month almost always failed one last month too. So the gap between 92% and 100% is not random noise. It is a short, stable list of devices, and those are precisely the machines that stay exposed to known vulnerabilities the longest.</p>\n    <p>The mistake is treating that list as a queue of identical chores. Reboot it, run the troubleshooter, watch it fail again in four weeks. The faster route is to work out which of the five causes each machine actually has, because the causes announce themselves quite clearly once you know the tells.</p>\n\n    <h2>How to tell which cause you actually have</h2>\n    <p>Start with the pattern, not the machine. <strong>One device failing repeatedly</strong> points to a local cause: disk, reboot state or corruption. <strong>Many devices failing at once</strong> points to infrastructure: they are all being told the same wrong thing. <strong>No errors at all, just staleness</strong>, points to machines that were never switched on when it mattered.</p>\n    <ul>\n      <li><strong>Full disk.</strong> The tell is error <code>0x80070070</code>, or simply a system drive with under 15 GB free. A cumulative update needs room to download, unpack and install, roughly triple its own size. Classic on older 128 GB SSDs.</li>\n      <li><strong>Pending reboot loop.</strong> The update reports installed, but compliance never changes and the device sits at \"restart required\" for weeks. Usual suspects: laptops that sleep every night and are never actually restarted.</li>\n      <li><strong>Corrupted update components.</strong> The same machine throws the same error code month after month, often <code>0x80073712</code> or another component-store error. The troubleshooter fixes it once, then it breaks again.</li>\n      <li><strong>Pointed at a server that no longer answers.</strong> Many machines \"checking for updates\" that never find any, often with <code>0x8024402f</code> or similar. Almost always a leftover WSUS address or update policy from a previous management tool, still set after the server was retired.</li>\n      <li><strong>Never online during the window.</strong> No errors anywhere. The machine is just behind, because it belongs to someone part-time, on leave, or it lives in a cupboard. Failure and absence look identical on a lazy report.</li>\n    </ul>\n\n    <h2>What to do about each</h2>\n    <p><strong>Full disk:</strong> clear space with Disk Cleanup or Storage Sense and remove old profiles, but be honest when the drive is simply too small. That machine is a hardware decision, not a patching one, and belongs in your <a class=\"inline\" href=\"/blog/hardware-refresh-cycle\">refresh plan</a>.</p>\n    <p><strong>Reboot loop:</strong> enforce restarts with a visible warning and a deadline rather than hoping users oblige. Then change what you measure: \"installed\" is not the finish line, \"installed and restarted\" is.</p>\n    <p><strong>Corruption:</strong> stop the update service, clear the <code>SoftwareDistribution</code> folder, then run <code>DISM /Online /Cleanup-Image /RestoreHealth</code> and <code>sfc /scannow</code>. If the same machine needs this twice, stop nursing it and rebuild it. A morning of reimaging is cheaper than a year of monthly surgery.</p>\n    <p><strong>Stale update server:</strong> check for a configured WSUS address in policy or the registry and remove it, so devices talk to Microsoft's servers directly. This one fix regularly resurrects whole fleets, especially after leaving an old RMM.</p>\n    <p><strong>Absent machines:</strong> decide deliberately. Wake devices for a maintenance window, or accept the lag but make reporting distinguish \"failed\" from \"not seen for 30 days\". Both need attention, but not the same attention.</p>\n    <p>One caution once the estate looks green: Windows updates are the visible half of patching. The browsers and readers that attackers actually target update <a class=\"inline\" href=\"/blog/third-party-patching-for-msps\">outside Windows Update entirely</a>, and deserve the same scrutiny.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Diagnosis is quick when the tells are in front of you, and tedious when each one means remoting into a machine. The Helios agent reports per-device patch state, disk space, pending reboots and last-seen times side by side, so the failed column separates into its real causes at a glance. When a device needs a closer look, Helio can investigate it and pull the update history and error codes for you, rather than you finding a login window.</p>"},{"slug":"msp-client-offboarding-checklist","title":"MSP client offboarding checklist: how to lose a client well","meta_description":"An MSP client offboarding checklist that protects both sides: notice and scope, inventory and credential handover, clean access removal and a final sign-off.","published_at":"2026-07-29T00:00:00","body_html":"<p class=\"lede\">Every MSP has an onboarding process. Far fewer have an offboarding one, because nobody enjoys planning for the day a client leaves. But clients do leave: they get acquired, they take IT in-house, they chase a cheaper contract. An MSP client offboarding checklist is what separates a professional exit that earns referrals from a messy one that ends in disputed invoices and awkward security questions. Here is ours, item by item, with what each step is actually protecting you from.</p>\n\n    <h2>What an MSP client offboarding checklist has to cover</h2>\n    <p>Offboarding is three things at once: a security event, because your access to someone else's estate is ending; a legal event, because contracts, data protection and licence ownership all come due; and a reputation event, because the way you leave is the story that client tells their peers for years. The <a class=\"inline\" href=\"/blog/msp-client-onboarding-checklist\">onboarding checklist</a> you run on day one deserves a mirror image at the end, and it splits naturally into four phases: agree the exit, hand over the estate, remove yourself, and close out cleanly.</p>\n\n    <h2>Phase one: agree the exit in writing</h2>\n    <p><strong>1. Re-read the contract before you respond to notice.</strong> Your agreement almost certainly specifies a notice period, what offboarding assistance you owe, whether you can charge for it, and how long you must retain data. Check it before you promise anything. Skip this and you end up negotiating terms mid-transition, doing weeks of handover work for free, or holding data longer or shorter than you agreed to, and every one of those becomes a dispute at exactly the moment goodwill is lowest.</p>\n    <p><strong>2. Fix a cutover date and name a contact on each side.</strong> A transition needs a single agreed end date, usually 30 to 60 days out, and one named technical contact at your firm, at the client, and at the incoming provider if there is one. Skip this and the offboarding drifts for months: you keep responding to tickets you are no longer paid for, nobody knows who owns an outage in week seven, and you carry liability without revenue.</p>\n\n    <h2>Phase two: hand over the estate, not a spreadsheet</h2>\n    <p><strong>3. Export a complete asset inventory.</strong> The successor needs every device, server and network appliance you managed, with OS versions, warranty status, patch state and assigned users, as of the cutover date. Skip it and every unknown machine discovered over the next year gets blamed on you, fairly or not. A dated, complete export is your evidence of what good order you left things in.</p>\n    <p><strong>4. Transfer documentation, credentials and ownership.</strong> Admin credentials for every in-scope system, network diagrams, the domain registrar login, DNS, Microsoft 365 admin roles, and any licences or vendor accounts registered in your name rather than the client's. Transfer ownership, not just passwords. Skipped, this is the classic horror story: eighteen months later the client's domain lapses or a certificate expires, the renewal notices are going to your billing inbox, and their website is down because of an account nobody remembered you owned.</p>\n\n    <h2>Phase three: remove yourself completely</h2>\n    <p><strong>5. Uninstall your agents and revoke your access.</strong> RMM agents, remote access tools, your technicians' admin accounts, API keys, delegated admin relationships. All of it, on the cutover date, and keep a record of the removal. Skip it and you retain silent access to an estate you no longer manage, which means that if anything goes wrong there, ever, you are a suspect with no contract protecting you.</p>\n    <p><strong>6. Check the places access hides.</strong> Break-glass accounts, your phone number on MFA resets, firewall rules allowing your network in, conditional access exclusions, backup consoles, and mail forwarding to your service desk. These outlive the obvious accounts because nobody documented them. Skipped, they surface in the client's next security audit as unexplained third-party access, and the finding has your company's name on it.</p>\n\n    <h2>Phase four: close out data and the relationship</h2>\n    <p><strong>7. Return, then delete, client data on a schedule you can evidence.</strong> Hand over ticket history, backups and stored files, then delete your copies once the agreed retention period ends, while holding anything a regulator or the contract obliges you to keep. Skip the deletion and you are a data protection liability holding personal data with no lawful basis; delete too eagerly and you may destroy something the client was legally required to retain.</p>\n    <p><strong>8. Write a final handover summary and get sign-off.</strong> One document: what was transferred, what was removed, what was deleted and when, and anything still outstanding. Ask the client to sign it. Skipped, you have no proof the exit was clean if questions come later, and you lose the closing conversation, which is where a well-run exit turns into a future referral. Clients who leave well come back, and they talk.</p>\n\n    <div class=\"callout\"><p><strong>Keep the tone boring.</strong> However the relationship ended, run the exit as if the client will be reading your emails aloud to your next prospect. Sometimes they effectively are: in most markets, the person who signed the cancellation letter shares a room with your future customers twice a year.</p></div>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Most of this checklist is discipline, not tooling, but the tooling decides how painful it is. Because Helios keeps a live per-client inventory of devices, patch state, protection status and tickets, the handover exports in phase two are a few clicks rather than a week of archaeology, ticket history can be pulled through the API, and removing a client is a clean, logged operation rather than a hunt through every endpoint. The same records that power a good <a class=\"inline\" href=\"/blog/msp-qbr-guide\">QBR</a> are the ones that make a good exit.</p>"},{"slug":"one-person-it-department","title":"One person IT department: how to make it work","meta_description":"Most one person IT departments fail by copying how big teams work. The case for running IT as a system: standardise, automate, and staff for the residue.","published_at":"2026-07-28T00:00:00","body_html":"<p class=\"lede\">If you are a one person IT department for a company of a hundred people, the industry has plenty of opinions about you, most of them written by firms that would like to replace you. Here is a different claim, and this whole article is a defence of it: one person can run a modern estate properly, but only by refusing to run it the way a big team does, and by treating every manual task as a design fault.</p>\n\n    <h2>The claim, stated properly</h2>\n    <p>One person can run IT for a hundred or a hundred and fifty users to a standard most larger departments would recognise as good. Patched machines, backups that have actually been restored from, MFA everywhere, tickets answered within a business day, an asset register that reflects reality. None of that requires a team. It requires a person who refuses to do repeatable work by hand.</p>\n    <p>The distinction matters because the common failure of solo IT is not competence. People who run IT alone are usually generalists of a rare kind, comfortable across networking, endpoints, Microsoft 365 and whatever the finance system throws at them. When they fail, it is because they run a department of one as if it were a scaled down department of ten: the same interrupt-driven queue, the same hand-built machines, the same ad hoc fixes, with a tenth of the hands.</p>\n    <p>A big team survives that model because it has slack. Someone can go deep on a project while someone else absorbs the interrupts. A department of one has no slack to hide in, so the model itself has to change. The job stops being fixing things and becomes designing an estate that mostly does not need fixing. Every hour spent on a task a machine could do is not just an hour lost. It is an hour that will be lost again next week, and every week after that.</p>\n\n    <h2>The maths that breaks the heroic model</h2>\n    <p>Common staffing benchmarks put one IT person per sixty to a hundred users, depending on how messy the estate is. Read that as a warning, not a target. The benchmark describes how much manual work a conventional estate generates. It says nothing about how much of that work needs to exist.</p>\n    <p>Look at what actually fills a solo operator's week. Password resets that self-service could handle. Patch runs babysat by hand. A monitoring inbox where the tenth false alarm of the morning gets deleted on sight, which is exactly the failure we described in <a class=\"inline\" href=\"/blog/msp-alert-fatigue\">our piece on alert fatigue</a>. A new starter set up from memory because the checklist lives in your head. None of this is the job. It is the interest payment on an estate that was never systematised.</p>\n    <p>Interrupt-driven work is also dearer than it looks, because every context switch has a re-entry cost and a department of one absorbs every interrupt personally. The queueing behaviour is the cruel part: when a single server runs close to capacity, wait times do not grow gradually, they explode. That is why solo IT feels fine at sixty percent load and catastrophic at ninety, and why \"busy but coping\" is usually the last stop before drowning.</p>\n    <p>Manual work scales linearly with headcount supported. Systems do not. That asymmetry is the entire argument, and everything below is an objection to it, taken seriously.</p>\n\n    <h2>Objection one: \"this is really an understaffing problem\"</h2>\n    <p>Sometimes it is, and it is worth being honest about the floor. If the business genuinely needs round-the-clock cover, carries heavy compliance obligations, or runs several sites that expect someone on the ground, one person cannot be the whole answer, and pretending otherwise just burns that person out. The argument is not that solo IT always works. It is that headcount is the wrong first fix.</p>\n    <p>Most one person IT departments that are drowning are not drowning in irreducible work. They are drowning in work that should not exist: resets without self-service, machines patched by hand, alerts nobody tuned, knowledge that lives in one head. Hiring into that fixes nothing. A second person doubles your capacity to do unnecessary work, and within a year you have a two person department that is also drowning, plus a salary. The order of operations matters: eliminate the work, automate what remains, then staff for the residue. If the residue still justifies a hire, hire with confidence, because the new person inherits a system instead of a pile.</p>\n    <div class=\"callout\"><p><strong>The residue test:</strong> write down everything you did last month, then cross off every task a system could have done: resets, patch runs, routine checks, chasing backup status. What is left is the real job. Staff for that, not for the pile.</p></div>\n\n    <h2>Objection two: \"a one person IT department has no time to standardise\"</h2>\n    <p>This objection gets the economics exactly backwards. Standardisation is not a luxury you buy once the fires are out. Variety is why the fires keep starting.</p>\n    <p>Every non-standard thing in an estate is a compounding cost. A snowflake laptop build means every incident on that machine starts with archaeology. Three PDF editors instead of one means three things to patch, licence and troubleshoot. Local admin rights granted \"temporarily\" two years ago are a malware surface nobody remembers approving. A big department can absorb variety with brute force. A department of one cannot, so it has to refuse the variety instead.</p>\n    <p>In practice that means one hardware standard per role on a predictable refresh cycle, one build deployed the same way every time, one approved application per job, and a default answer of no to one-off requests, with an honest path for the rare genuine exception. None of this needs a committee, and that is the solo operator's quiet advantage: there is no committee. You can decide the standard on Tuesday and enforce it on Wednesday. Say yes to every one-off instead, and you will never notice the moment the estate became unmanageable, because it happens one reasonable exception at a time.</p>\n\n    <h2>Objection three: \"unattended automation is too dangerous\"</h2>\n    <p>The fear is reasonable on its face. If an automated fix misfires at two in the morning, there is no colleague to catch it. But this objection compares automation to a perfect operator who watches everything, and that operator does not exist. The real comparison is automation with guardrails against an estate drifting unwatched while its one human is heads-down in a ticket.</p>\n    <p>The unattended risk you already run is far larger than the one you are worried about. Machines quietly months behind on patches. An endpoint whose antivirus lapsed after an update. A backup job that has failed nightly for weeks with nobody reading the report. Manual operation does not prevent any of this. It causes it, because one person's attention cannot be everywhere, and each of these fails silently.</p>\n    <p>So sequence the automation rather than refusing it. First, visibility: an inventory and monitoring baseline that tells the truth, so \"I did not know\" stops being a possible sentence. Second, alerting you actually trust, tuned hard enough that every alert means something. Only then remediation, and even then with guardrails: start with actions that are safe to repeat, roll changes to a small ring of machines before the fleet, and log everything so you can read at nine in the morning what happened at two. Built this way, automation is not a junior you cannot supervise. It is the only colleague a department of one can afford.</p>\n\n    <h2>Objection four: \"you cannot do security alone\"</h2>\n    <p>You cannot run a security operations centre alone, and you should not try. But most breaches of smaller organisations do not need a SOC to prevent, because they are not sophisticated. They walk in through an unpatched browser, a mailbox without MFA, a forwarding rule nobody noticed, or a machine whose protection silently stopped.</p>\n    <p>That list is tractable for one person precisely because it can be systematised. MFA enforced everywhere and legacy authentication off. Patching that covers third-party applications, not just Windows. A standing check that every endpoint really is protected, because the gap between \"we bought antivirus\" and \"it is running and current on every machine\" is where incidents live. Backups that get restore-tested rather than admired for their green ticks. And a hardened Microsoft 365 baseline, which is a checklist rather than a discipline: <a class=\"inline\" href=\"/blog/microsoft-365-security-checklist\">our ten-setting Microsoft 365 checklist</a> is a workable start.</p>\n    <p>Hold that line and you are ahead of plenty of organisations with far larger teams, because these controls stop the attacks that actually arrive. Where outside help genuinely earns its keep is at the edges: an incident response retainer for the bad day, and after-hours cover if the business truly needs it. Buy those as insurance, not as a substitute for basics that only you are positioned to enforce.</p>\n\n    <h2>Objection five: \"what happens when you go on holiday?\"</h2>\n    <p>The bus factor is the most honest objection, and it is aimed at the wrong target. The risk was never that IT is one person. The risk is that the estate exists only in that person's head. A heroic department of one really is dangerous to its business: undocumented, irreplaceable and one resignation away from chaos.</p>\n    <p>A systems-run department of one is the opposite, because the system is the documentation. The asset register is live rather than a spreadsheet from March. Patch state, protection state and backup state are visible on a screen anyone can read. The runbooks are not prose in a wiki, they are the automations themselves, and they keep running whether or not you are on a beach. What remains to write down is small: escalation contacts, the credentials process, and the reasoning behind your standards. Add one arrangement for cover, a retainer with a local provider or a reciprocal agreement with another solo operator, and the holiday question is answered better than many mid-sized departments answer it.</p>\n    <p>Then hold yourself to the test this whole argument implies. Each month, ask what share of issues resolved without you touching them, and how much of your time went on work only you could have done. A one person IT department run on heroics watches those numbers stand still. One run as a system watches the first climb and the second shrink toward the work that genuinely deserves a human.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios was built on the premise that leverage should come from the platform, not from headcount. One lightweight agent provides the live inventory, patch state including third-party applications, protection gap detection and backup monitoring this argument depends on. The service desk and client portal give the people you support a proper front door, and Helio, the AI layer, triages tickets, investigates devices and resolves routine failures on its own, with every action logged for you to read the next morning. A one person IT department is not an edge case for us. It is close to the design brief.</p>"},{"slug":"hardware-refresh-cycle","title":"Hardware refresh cycles: how often to replace laptops, desktops and servers","meta_description":"How often should you replace laptops, desktops and servers? Practical hardware refresh cycle guidance, replacement signals, and how to budget a rolling refresh.","published_at":"2026-07-27T00:00:00","body_html":"<p class=\"lede\">Every IT budget conversation eventually arrives at the same question: how long should this hardware last? A sensible hardware refresh cycle is the difference between replacing machines on your terms and replacing them in a panic when they fail. This guide covers realistic lifespans by device class, the signals that matter more than age, and how to budget a rolling refresh, whether you look after one company's estate or a dozen clients' worth.</p>\n\n    <h2>What a hardware refresh cycle actually is</h2>\n    <p>A hardware refresh cycle is a standing policy that says how long each class of device stays in service before it is replaced, and a plan that spreads those replacements across budget years. That is all. It is not an enterprise asset management programme, and you do not need one to have a refresh cycle. You need an accurate inventory, an agreed lifespan per device class, and a spreadsheet or a platform that can tell you what falls due next year.</p>\n    <p>The alternative, running everything until it dies, feels thrifty and is usually the most expensive option available. An aged machine does not fail politely at a convenient moment. It fails on a Tuesday morning with a deadline attached, and the true cost includes the emergency purchase at whatever price is available that week, the hours rebuilding a user's environment from nothing, and the productivity lost while all of that happens. Planned replacement moves the same spend to a moment you chose, at a price you negotiated, with the user's data migrated calmly in advance.</p>\n\n    <h2>How often should you replace laptops, desktops and servers?</h2>\n    <p>There is no single number, but there are defensible ranges that most of the industry has converged on:</p>\n    <ul>\n      <li><strong>Laptops: 3 to 4 years.</strong> They travel, they get dropped, batteries degrade and hinges wear. By year four, repair cost and lost productivity usually overtake replacement cost. Machines doing heavy work, development, CAD, video, sit at the short end.</li>\n      <li><strong>Desktops: 4 to 5 years.</strong> No battery, no travel, easier to upgrade. A desktop that was decently specified at purchase can do five years without the user suffering for it.</li>\n      <li><strong>Servers: 5 to 7 years.</strong> The practical limit is usually vendor support and parts availability rather than the hardware itself. Running production workloads on a server the vendor no longer supports is a risk decision, not a savings decision.</li>\n      <li><strong>Network kit and firewalls: 5 to 8 years.</strong> Switches are long-lived, but firewalls are only as good as their security updates. The day a firewall leaves vendor support is the day it becomes a liability, whatever its uptime says.</li>\n    </ul>\n    <p>Align the lifespan with the warranty where you can. Buying laptops with a three-year warranty and running them for four means the final year is uninsured, which is exactly when failures cluster. Either buy the extended warranty or shorten the cycle to match.</p>\n\n    <h2>The Windows 11 long tail is forcing the issue</h2>\n    <p>Plenty of estates are having this conversation now whether they planned to or not. Windows 10 left support in October 2025, and Windows 11's hardware requirements, TPM 2.0 and a roughly 2018-or-newer processor, mean a real share of older machines cannot simply be upgraded in place. Those devices are not just old, they are stuck: still working, still logged in, and no longer receiving free security updates.</p>\n    <div class=\"callout\"><p><strong>Deadline worth diarising:</strong> the Extended Security Updates bridge is temporary. Consumer ESU ends in October 2026, and paid commercial ESU roughly doubles in price each year it is renewed. If you still have Windows 10 machines in service, the refresh plan for them needs a date on it this quarter, not a vague intention.</p></div>\n    <p>The useful lesson generalises beyond this one migration: operating system support windows are part of hardware lifespan. A machine that cannot run the next supported OS is end-of-life on a schedule someone else set, and your refresh cycle should see that coming years out, because the ship dates are published years in advance.</p>\n\n    <h2>Age is a proxy: the signals that matter more</h2>\n    <p>Calendar age is the planning number, but it is a proxy for the things you actually care about. A four-year policy applied blindly replaces some machines that are fine and misses some that are quietly ruining someone's working day. The better refresh decisions weigh signals like these:</p>\n    <ul>\n      <li><strong>Disk health.</strong> SMART warnings and reallocated sectors are the closest thing hardware gives you to advance notice. A machine with a degrading disk goes to the top of the list regardless of age.</li>\n      <li><strong>Battery wear.</strong> A laptop that no longer survives a meeting without a charger has effectively become a desktop, and its user knows it.</li>\n      <li><strong>Ticket history.</strong> Three tickets in six months from the same device is a pattern. The support time already spent on it is usually a meaningful fraction of the replacement cost.</li>\n      <li><strong>Performance drag.</strong> Boot times, memory pressure and CPU saturation during ordinary work. Users rarely report slowness, they just absorb it, so measure it rather than waiting to be told.</li>\n      <li><strong>Support status.</strong> Warranty expiry, OS eligibility and vendor end-of-support dates, as above.</li>\n    </ul>\n    <p>The practical approach is a policy with an override in each direction: devices are scheduled by age, promoted early when the signals are bad, and occasionally kept a little longer when a healthy machine is doing light duty for an undemanding user.</p>\n\n    <h2>Rolling refresh beats big bang: the budget maths</h2>\n    <p>The other half of a refresh cycle is how the spend lands, and the answer is almost always: spread it. Take an estate of 100 laptops on a four-year cycle. Replacing them as one project means a large procurement roughly every four years, a bruising budget line, a deployment crunch, and then four quiet years in which the entire fleet ages in lockstep towards the next cliff.</p>\n    <p>The rolling alternative replaces a quarter of the fleet every year: 25 machines, every year, forever. The total spend over four years is identical, but everything else improves. The budget line is flat and predictable, which finance departments and clients both prefer. Deployment becomes routine rather than a project. The fleet always contains a spread of ages, so no single OS deadline or bad batch can strand half your machines at once. And each year's purchase benefits from that year's price and performance rather than locking the whole estate to one moment in time.</p>\n    <p>The same logic scales down. An estate of 20 laptops on a four-year cycle is five machines a year, which is a purchase order, not a project. If you are starting from an estate that has never had a refresh plan, begin with the oldest quartile this year and the cycle establishes itself automatically.</p>\n\n    <h2>Building a refresh plan you can defend</h2>\n    <p>A workable plan takes an afternoon if the inventory exists, and it looks like this:</p>\n    <ol>\n      <li><strong>Get the inventory straight.</strong> Every device, with model, purchase or first-seen date, warranty status, OS eligibility and assigned user. If devices are appearing that nobody procured, fix discovery first: a refresh plan built on a partial inventory is fiction. This is the same discovery discipline that anchors a good <a class=\"inline\" href=\"/blog/msp-client-onboarding-checklist\">client onboarding</a>, and it pays off here for years.</li>\n      <li><strong>Assign a lifespan per class</strong>, using the ranges above adjusted for how hard your machines actually work.</li>\n      <li><strong>Generate the due list.</strong> Sort by age against lifespan and you have next year's replacement list and a defensible number to put in the budget. Multiply out the following two years while you are there: a three-year forward view is what turns hardware from a surprise into a line item.</li>\n      <li><strong>Apply the overrides.</strong> Promote the machines with bad signals, demote the healthy outliers, and note why, so the plan survives scrutiny.</li>\n      <li><strong>Review annually.</strong> Lifespans drift, requirements change, and last year's exceptions should not silently become policy.</li>\n    </ol>\n    <p>If you run an MSP, this plan is also one of the easiest pieces of proactive value you can show. A three-year hardware forecast per client, presented once a year with prices attached, answers the \"what are we paying you for\" question better than any ticket count, and it turns emergency hardware purchases, the thing clients remember and resent, into planned ones.</p>\n\n    <h2>Retirement is part of the cycle</h2>\n    <p>A refresh plan that stops at \"new laptop delivered\" is only half finished. The outgoing device still holds company data and still shows up in licence counts, and old machines have a way of living in cupboards as unmanaged, unpatched spares that resurface on the network a year later. Close the loop deliberately: wipe the disk to a documented standard or destroy it, keep a record of the disposal for compliance, recycle through a certified e-waste route, and mark the asset retired in the inventory the same day. A short written checklist here is worth more than good intentions, because disposal is the step nobody is ever chased for.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Almost everything above depends on data you should not have to collect by hand. Helios already maintains the live inventory its agent sees, hardware model, age, disk health, OS version and warranty-relevant details, alongside each device's ticket history, so the asset lifecycle view can surface what is due, what is degrading early and what is quietly costing you support time. That evidence is what turns a hardware refresh cycle from a guess into a schedule, and because the same data sits behind <a class=\"inline\" href=\"/blog/third-party-patching-for-msps\">patching</a> and monitoring, the machine you are about to replace and the machine you are about to patch are finally the same record.</p>"},{"slug":"microsoft-365-security-checklist","title":"Microsoft 365 security checklist: the ten settings to fix first","meta_description":"A practical Microsoft 365 security checklist: the ten settings to fix first, from MFA and legacy authentication to mail forwarding and offboarding.","published_at":"2026-07-26T00:00:00","body_html":"<p class=\"lede\">A new Microsoft 365 tenant is configured for adoption, not for safety. Microsoft's defaults are chosen so that nothing gets in the way on day one, which means legacy protocols linger, sharing is open, and admin rights sprawl. This Microsoft 365 security checklist covers the ten settings to fix first, whether you look after one tenant for your own organisation or dozens for clients.</p>\n\n    <h2>Why the defaults are not enough</h2>\n    <p>Most Microsoft 365 breaches do not involve anything clever. They are password sprays against accounts without MFA, sign-ins through legacy protocols that cannot enforce it, a compromised mailbox quietly forwarding invoices to an attacker, or an ex-employee account that stayed live for six months. Every one of those is closed by configuration, not by buying anything.</p>\n    <p>Microsoft has been tightening the baseline, and security defaults now come switched on for new tenants, but \"on by default\" is not the same as \"configured for your organisation\". Defaults do not know which of your accounts are admins in practice, what your sharing posture should be, or who left last month. The checklist below is deliberately short and ordered by impact. Most items cost nothing beyond the licences you already have.</p>\n\n    <h2>The Microsoft 365 security checklist</h2>\n    <ol>\n      <li><strong>Enforce MFA for every account, no exceptions.</strong> Not just admins, and not \"registered but optional\". A single unprotected mailbox is a foothold, and attackers do not care whose it is. Conditional Access is the cleaner way to enforce it if your licensing includes it; security defaults are the fallback if not.</li>\n      <li><strong>Give admins stronger authentication and separate accounts.</strong> Admin roles should sit on dedicated accounts that are not used for daily email, protected by phishing-resistant methods such as passkeys or FIDO2 keys rather than SMS codes. An admin who reads mail with the same account that holds Global Administrator is one convincing phish away from handing over the tenant.</li>\n      <li><strong>Cut the Global Administrator list and create a break-glass account.</strong> Two to four global admins is plenty for almost any organisation. Everyone else gets a scoped role: Exchange admin, Helpdesk admin, User admin. Then create one emergency access account, excluded from Conditional Access, with a very long password stored offline, so a bad policy change cannot lock you out of your own tenant.</li>\n      <li><strong>Block legacy authentication.</strong> Protocols like IMAP, POP and SMTP basic auth cannot do MFA, which makes them the open window next to your locked front door. Check the sign-in logs first for anything still using them, fix or retire it, then block the lot with Conditional Access.</li>\n      <li><strong>Confirm unified audit logging is on and actually recording.</strong> When something does go wrong, the audit log is the difference between an investigation and a guess. It is enabled by default on current tenants, but verify it, and run a test search so the first time you use it is not during an incident.</li>\n      <li><strong>Turn on the preset email security policies.</strong> Exchange Online Protection and, if licensed, Defender for Office 365 ship with Standard and Strict presets covering anti-phishing, anti-spoofing, Safe Links and Safe Attachments. The presets are better than most hand-rolled policies and they update as threats change. Apply Standard broadly and Strict to finance and executives.</li>\n      <li><strong>Rein in external sharing.</strong> The default SharePoint and OneDrive posture allows broad external sharing. Set the tenant default to \"new and existing guests\" or tighter, put expiry on guest links, and review existing guest accounts quarterly. Most organisations that do this for the first time are surprised by what has been shared and forgotten.</li>\n      <li><strong>Block auto-forwarding and watch for new inbox rules.</strong> The classic sign of a compromised mailbox is a rule that forwards or deletes mail so the owner never sees the replies. Disable automatic external forwarding at the transport level and alert on new forwarding rules. This one setting has quietly defeated a large share of business email compromise attempts.</li>\n      <li><strong>Tie sign-ins to devices you trust.</strong> Where licensing allows, require compliant or Entra-joined devices for access to company data, or at minimum block sign-ins from countries you never operate in. Identity controls work best when the endpoint itself is known and <a class=\"inline\" href=\"/blog/third-party-patching-for-msps\">properly patched</a>, so treat device health and tenant hardening as one exercise.</li>\n      <li><strong>Make offboarding immediate and complete.</strong> Disable the account, revoke active sessions, remove it from groups, convert or archive the mailbox, and reclaim the licence, on the day the person leaves. Stale accounts are the cheapest attack surface there is, and they also cost you real money in licences.</li>\n    </ol>\n\n    <div class=\"callout\"><p><strong>Do not over-harden in one afternoon.</strong> The most common failure mode is enthusiasm: someone applies every control on a Friday, Monday brings a flood of locked-out users, and the changes get rolled back in a panic, often further back than where you started. Roll out in stages, use report-only mode for Conditional Access policies first, and communicate before you enforce.</p></div>\n\n    <h2>Settings drift, so a checklist is not a one-off</h2>\n    <p>A tenant hardened in January is not hardened in July. New starters join outside the MFA policy, someone grants a contractor Global Administrator \"temporarily\", a vendor asks for an app consent that never gets reviewed, and a sharing exception made for one project becomes permanent. None of these announce themselves.</p>\n    <p>The fix is the same discipline that applies to backups: checking once proves nothing, so verify continuously. A green tick you saw six months ago <a class=\"inline\" href=\"/blog/backup-monitoring-for-msps\">is not evidence of anything today</a>. Put a quarterly review in the calendar, track Secure Score as a trend rather than a trophy, and treat any drop as a question to answer, not a number to ignore.</p>\n\n    <h2>Making it stick as a small team</h2>\n    <p>If you are a one or two person team, or an MSP juggling many tenants, the checklist only works if it is cheap to repeat. Three habits keep it that way:</p>\n    <ul>\n      <li><strong>Write down your intended state.</strong> One page per tenant: who the global admins are, what the sharing default is, which Conditional Access policies exist and why. Drift is invisible without a baseline to compare against.</li>\n      <li><strong>Alert on the changes that matter.</strong> New admin role assignments, new forwarding rules, and risky sign-ins deserve an alert. Everything else can wait for the quarterly review.</li>\n      <li><strong>Automate the checking, not just the fixing.</strong> The scarce resource is attention. Anything that turns \"log in and look\" into \"get told when it changes\" pays for itself within a month.</li>\n    </ul>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios connects to Microsoft 365 alongside the devices, patching, antivirus and backup state it already monitors, so tenant posture sits in the same view as the rest of the estate instead of in a portal you remember to check. Helio, the AI layer, uses that context when it triages tickets and investigates devices: a sign-in problem looks very different when the platform can see both the endpoint and the identity side. The checklist above is exactly the kind of routine verification we think platforms should carry for you.</p>"},{"slug":"msp-qbr-guide","title":"MSP QBR guide: quarterly business reviews clients actually value","meta_description":"What to put in an MSP QBR, what to leave out, and how to prepare one in under an hour: a 45-minute agenda, the metrics that matter, and the mistakes that lose clients.","published_at":"2026-07-25T00:00:00","body_html":"<p class=\"lede\">The MSP QBR is the meeting everyone agrees is important and almost nobody runs well. Done badly, a quarterly business review is an hour of ticket counts a client politely endures. Done well, it is the single strongest retention tool an MSP has: the one moment a quarter when the client sees what they are paying for and what should happen next.</p>\n\n    <h2>Why QBRs quietly fall off the calendar</h2>\n    <p>Most MSPs do not skip QBRs because they doubt the value. They skip them because preparation is expensive. Pulling ticket stats from the PSA, patch numbers from the RMM, backup results from another console, then assembling it all into a deck takes half a day per client. Multiply that by thirty clients and a quarterly cadence becomes a full-time job nobody has.</p>\n    <p>So the meetings drift to twice a year, then to \"when the contract is up for renewal\", which is exactly the wrong time. By renewal, the client has already formed a view of your value based on the only signal they have seen all year: whether things broke. If nothing broke, they wonder what they are paying for. If something did, that is what they remember. The QBR exists to replace that lopsided impression with evidence.</p>\n\n    <h2>What an MSP QBR is actually for</h2>\n    <p>A QBR is not a service report read aloud. It has three jobs, in this order:</p>\n    <ol>\n      <li><strong>Prove the value of the last quarter.</strong> Not with raw activity, but with outcomes: risks reduced, incidents prevented, commitments met.</li>\n      <li><strong>Surface decisions the client needs to make.</strong> Ageing hardware, unsupported software, security gaps you cannot close without their budget or sign-off. The QBR is where \"we recommended this in writing\" gets its date stamp.</li>\n      <li><strong>Agree what happens next quarter.</strong> Two or three concrete actions with owners. This is what separates a business review from a status update.</li>\n    </ol>\n    <p>Notice what is not on that list: selling. Upsell is a frequent by-product of a good quarterly business review, because gaps become visible and budgets get discussed. But the moment a client senses the meeting exists to sell them something, they stop attending, and the retention value dies with it.</p>\n\n    <h2>A 45-minute agenda that works</h2>\n    <p>Shorter is better. A tight 45 minutes with a decision-maker beats a rambling 90 with whoever could make it. A structure that holds up:</p>\n    <ul>\n      <li><strong>Five minutes: the quarter in three numbers.</strong> Lead with the headline posture, not the detail. Patch compliance, protection coverage, backup success. Green where it is green, honest where it is not.</li>\n      <li><strong>Ten minutes: service performance.</strong> Tickets by priority, response and resolution against targets, and any breaches with what changed as a result. If you have not defined those targets tightly, fix that first: our guide to <a class=\"inline\" href=\"/blog/msp-sla-response-times\">MSP SLA response times</a> covers what realistic ones look like.</li>\n      <li><strong>Ten minutes: risk and lifecycle.</strong> Devices leaving support, warranties expiring, operating systems ageing out, single points of failure. This is the section clients remember, because it is about their business, not your dashboard.</li>\n      <li><strong>Ten minutes: recommendations and decisions.</strong> Two or three items maximum, each with a cost, a risk statement and a requested decision. Ten recommendations is a list; three is a plan.</li>\n      <li><strong>Ten minutes: their agenda.</strong> Hiring plans, office moves, new software, worries. What you learn here fills the next quarter's roadmap and often surfaces work you would never have heard about otherwise.</li>\n    </ul>\n\n    <div class=\"callout\"><p><strong>One rule that improves every QBR:</strong> never present a number you cannot defend if the client asks \"compared to what?\" Every metric needs either a target, a previous quarter, or a peer baseline next to it. A lone \"97%\" means nothing; \"97%, up from 91% last quarter, against a 95% target\" is a story.</p></div>\n\n    <h2>The numbers worth showing, and the ones to leave out</h2>\n    <p>Clients do not care how busy you were. They care whether they are safer, faster and better prepared than last quarter. Metrics that earn their slide:</p>\n    <ul>\n      <li><strong>Patch compliance</strong> across the estate, including third-party applications, with the trend over the quarter.</li>\n      <li><strong>Endpoint protection coverage</strong>: how many devices are fully protected, and how many gaps were found and closed.</li>\n      <li><strong>Backup health</strong>: success rates, yes, but also restore tests performed. A green job history is not the same as a proven recovery, a distinction we cover in <a class=\"inline\" href=\"/blog/backup-monitoring-for-msps\">backup monitoring for MSPs</a>.</li>\n      <li><strong>Response and resolution against agreed targets</strong>, by priority, with breaches acknowledged rather than hidden.</li>\n      <li><strong>Asset lifecycle position</strong>: the count of devices past or approaching end of support, and the replacement plan.</li>\n    </ul>\n    <p>And the ones to drop: total tickets closed with no context, uptime percentages nobody disputed, screenshots of your monitoring console, and any metric whose only purpose is demonstrating effort. Effort is an input. QBRs are for outcomes.</p>\n\n    <h2>Preparing a QBR without losing half a day</h2>\n    <p>The preparation problem is real, and the fix is structural rather than heroic: the data should already be in one place. If patching, endpoint protection, backup status, tickets and asset inventory live in one platform, the evidence section of a QBR is an export, not an archaeology project. The half-day assembly job exists only where the tooling is fragmented.</p>\n    <p>Two habits make the rest cheap. First, keep a running \"QBR notes\" entry per client and add to it the moment something notable happens: a prevented incident, a recurring issue, a risk spotted. Reconstructing a quarter from memory is the slowest part of preparation. Second, standardise the deck. The same structure for every client every quarter means preparation becomes filling in numbers, and clients learn to read it at a glance.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios keeps the evidence a QBR needs in one place by design: patch compliance including third-party applications, endpoint protection gaps, Veeam and Acronis backup status, ticket performance and asset lifecycle all sit against each client already. Because Helio, the AI layer, investigates and resolves issues as the quarter runs, the record of what was prevented builds itself, which is precisely the material that proves value in the room. The half-day of assembly becomes minutes, which is the difference between QBRs that happen and QBRs that slip.</p>"},{"slug":"do-you-still-need-an-msp","title":"Do you still need an MSP? An honest answer from someone who runs one","meta_description":"For twenty years the answer was automatic. AI changed what a platform can do unaided, so here is an honest look at when you still need an MSP and when you do not.","published_at":"2026-07-25T00:00:00","body_html":"<p class=\"lede\">For most of two decades the answer was automatic. Too small for your own IT department, so you outsourced to a managed service provider. I want to question that default, which is an odd thing to do when I run an MSP myself. But the honest answer was never yes or no. It depends entirely on which of three things you were actually paying for, and one of those three has just changed underneath all of us.</p>\n\n    <h2>Serious tooling was gatekept, and not by technology</h2>\n    <p>Start with the thing nobody says out loud. For twenty years, the software that lets you run IT properly has been kept away from small organisations by pricing models rather than by any technical limit.</p>\n    <p>Per technician. Per endpoint. A sales call before you are allowed to see a number. Three year terms with the auto renewal buried in clause 14. The message was consistent: this level of control is for organisations with a procurement department, or for the MSPs who aggregate everybody else's demand. So if you were a two person IT team, a founder who knows their way around a server, or a junior engineer trying to do things properly, you had two realistic options. Pay the big vendors, or hold it together with duct tape and goodwill.</p>\n    <p>I think that world is ending, and I think it deserves to.</p>\n\n    <h2>What you were actually buying from an MSP</h2>\n    <p>Strip away the brochure language and an MSP contract has always bundled three quite different things:</p>\n    <ol>\n      <li><strong>Tooling.</strong> Monitoring, patching, a service desk, security visibility, remote access.</li>\n      <li><strong>Expertise.</strong> People who know what an alert means, which patch is safe to push on a Friday, why Outlook keeps crashing on one machine, and what to do at 2am.</li>\n      <li><strong>Accountability.</strong> A name on a contract. An SLA. Someone whose job it is when it breaks.</li>\n    </ol>\n    <p>The bundle made sense because the first two were inseparable. Tooling was close to useless without the expertise to act on it, and expertise did not scale without the tooling. Aggregating demand for both is exactly why MSPs exist, and small businesses were entirely rational to pay the margin for it.</p>\n\n    <h2>What changed is where the expertise sits</h2>\n    <p>The tooling got good enough to carry part of the expertise.</p>\n    <p>I do not mean a summary button in the corner of a dashboard. I mean software that reads the event log on the actual machine, works out why Outlook keeps crashing, writes the fix, asks permission, runs it and then verifies that it worked. That used to be a senior engineer's afternoon. It is now something a platform does while you are in a meeting. The same shift is happening to the ticket queue: the easy half of inbound requests can be answered before a human reads them.</p>\n    <p>The scarce resource was never really the software. It was the senior engineer hours needed to operate it. That is the constraint that has collapsed, and it is the constraint the old pricing was built on.</p>\n    <p>This does not make expertise worthless. It moves the line. A capable generalist with the right platform can now look after an estate that would have needed a small specialist team five years ago.</p>\n\n    <h2>Who genuinely does not need an MSP any more</h2>\n    <p>Be honest with yourself about which of these describes you:</p>\n    <ul>\n      <li><strong>You have one or two capable IT people</strong> who spend the week firefighting instead of improving anything. They do not lack skill. They lack leverage. A platform that investigates, fixes and triages is that leverage.</li>\n      <li><strong>You are technical enough and the estate is modest.</strong> Under a hundred devices, with a founder or ops lead who can hold it. The 2am problem is real but rare, and a platform that resolves what it can and only wakes you for the rest covers most of it.</li>\n      <li><strong>You are a startup MSP.</strong> You are the counter example that proves the point. The same shift that lets in-house teams serve themselves lets you take on clients from day one without the bench you cannot yet afford.</li>\n    </ul>\n\n    <h2>Who should absolutely still use an MSP</h2>\n    <p>This piece would be dishonest without this section, and since I run an MSP I know exactly where the line sits.</p>\n    <ul>\n      <li><strong>Nobody internal owns IT.</strong> A platform with no owner is a dashboard nobody opens. If no one in the building will genuinely hold it, pay someone whose job it is.</li>\n      <li><strong>Compliance carries the weight.</strong> If your insurer, regulator or largest customer expects a named accountable provider, then accountability is the product you are buying, not the tooling.</li>\n      <li><strong>You need hands.</strong> Hardware, offices, cabling, starter and leaver days. Software does not crawl under a desk.</li>\n      <li><strong>The estate is genuinely complex.</strong> Multiple sites, regulated data, legacy systems held together with string and institutional memory. Experience earns its margin here, easily.</li>\n    </ul>\n    <p>So the interesting conclusion is not that MSPs are finished. It is that the lazy default is finished. Any MSP whose fee was justified mainly by gatekeeping access to tooling is in trouble. Any MSP that brings judgement, accountability and hands on top of a modern platform is more valuable than it has ever been, because the platform takes away the toil that was compressing its margins in the first place.</p>\n\n    <h2>The test to apply to any platform, including mine</h2>\n    <p>If you are going to hold IT in your own hands, the platform is the hire. So interview it like one. I would hold any vendor to this, and that includes us:</p>\n    <ul>\n      <li>Published pricing, flat, with no per seat games</li>\n      <li>Every feature on every plan, so you are never upsold in the middle of an incident</li>\n      <li>AI that acts and shows its working on the machine, rather than a chat window bolted onto a dashboard</li>\n      <li>Guardrails you control: suggest only, approve first, or fully autonomous, set per site or per client</li>\n      <li>Monthly billing you can cancel from your own account page, without an email to anyone</li>\n      <li>Your data reachable over an API, so leaving is always genuinely possible</li>\n    </ul>\n    <div class=\"callout\">\n      <p>Any vendor that fails that last point is quietly telling you who holds the control. It is worth noticing before you sign, not after.</p>\n    </div>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Full disclosure, because it matters: I build one of these platforms, and I run my own MSP on it every day. So of course I would argue all this. Do not take my word for any of it.</p>\n    <p>Helios is my answer to the shift: monitoring, patching, security, Microsoft 365, backup monitoring, remote access and an AI service desk in one product, at <a class=\"inline\" href=\"/#pricing\">one flat price per organisation</a> with every feature on every plan. It is built multi-client for MSPs, and an in-house team simply runs its own organisation as the only client. The AI, Helio, <a class=\"inline\" href=\"/docs/helio-ai\">investigates alerts on the machine itself</a>, fixes what you allow it to fix, and answers the routine half of your users' tickets. We publish <a class=\"inline\" href=\"/compare/\">honest comparisons with the alternatives</a>, including the places they beat us, and a <a class=\"inline\" href=\"/trust\">plain English account of how we handle security</a>. Judge it against the list above rather than against my opinion.</p>\n    <p>The tools finally got good enough that you do not automatically need the big guys. What you need is the right platform, and the nerve to take control of it.</p>"},{"slug":"msp-sla-response-times","title":"MSP SLA response times: setting targets you can actually hit","meta_description":"Practical MSP SLA response times by priority, how to define response versus resolution, and how to set service desk targets your team can actually hit.","published_at":"2026-07-24T00:00:00","body_html":"<p class=\"lede\">Most MSP SLA response times were copied from someone else's contract. They looked reassuring in the proposal, nobody has measured them since, and the first person to read them closely is an unhappy client with a lawyer. This guide covers what response actually means, sensible targets by priority, and how to set numbers your service desk can hit every month, not just in a good week.</p>\n\n    <h2>Response is not resolution, and both need defining</h2>\n    <p>The single most common SLA failure is not a missed target. It is a contract where \"response\" was never defined, so the MSP and the client are measuring two different things in good faith.</p>\n    <p>Pin down three separate clocks:</p>\n    <ul>\n      <li><strong>First response:</strong> a qualified human has read the ticket, triaged it and told the client what happens next. An automated \"we got your email\" receipt does not count, and clients notice when it is passed off as one.</li>\n      <li><strong>Resolution:</strong> service is restored or a workaround is in place. Not \"root cause fully understood\", which can take much longer and belongs in a post-incident note, not the SLA clock.</li>\n      <li><strong>Updates:</strong> how often the client hears from you while a ticket is open. On a P1 that might be every 30 minutes. Silence during an outage does more contract damage than the outage itself.</li>\n    </ul>\n    <p>Write all three into the agreement, per priority level. If your PSA or service desk cannot report on all three, that is a tooling gap to fix before you sign anything.</p>\n\n    <h2>Typical MSP SLA response times by priority</h2>\n    <p>Numbers vary by market and price point, but most healthy MSP agreements land close to this shape:</p>\n    <ul>\n      <li><strong>P1, critical:</strong> a whole site or core system is down. First response in 15 to 30 minutes, work continues until service is restored, updates at least every 30 to 60 minutes.</li>\n      <li><strong>P2, major:</strong> a team or key function is degraded. First response within 1 to 2 hours, resolution targeted the same business day.</li>\n      <li><strong>P3, minor:</strong> one user affected, work continues with friction. First response within 4 business hours, resolution by the next business day.</li>\n      <li><strong>P4, request:</strong> new starters, changes, questions. First response within 1 business day, fulfilment within 3 to 5.</li>\n    </ul>\n    <p>Two caveats before you lift these into a contract. First, they assume business-hours cover; if you sell 24/7, the P1 line is the expensive one, so price it deliberately. Second, these are ceilings, not ambitions. If your desk routinely responds to P3s in 20 minutes, keep the 4-hour promise anyway. The gap between what you promise and what you deliver is your margin for the bad week, and every desk has bad weeks.</p>\n\n    <div class=\"callout\"><p><strong>Rule of thumb:</strong> set the contractual target at roughly double your real-world average. If you consistently hit first response on P2s in 25 minutes, promise an hour. You will beat the SLA month after month, and the report that proves it becomes a retention tool rather than a liability.</p></div>\n\n    <h2>Priority must come from impact, not volume</h2>\n    <p>An SLA is only as good as the triage behind it. If priority is set by whoever shouts loudest, every ticket from the shouty client becomes a P1, your team burns out chasing false urgency, and the genuinely critical ticket from a quiet client waits.</p>\n    <p>The fix is a simple, written impact and urgency matrix: how many people are affected, and can they work at all? A payroll server down on the 28th of the month is a P1 even if the ticket arrived politely. A director's second monitor is a P3 even if it arrived in capitals. Put the matrix in the client agreement too, so priority is a shared definition rather than a negotiation on every ticket.</p>\n    <p>Triage quality is also where alert-driven tickets earn their keep or destroy it. If your monitoring floods the desk with noise, technicians stop trusting priorities altogether, and the SLA clock runs on tickets nobody should be seeing. We covered how to fix that at the source in <a class=\"inline\" href=\"/blog/msp-alert-fatigue\">our guide to cutting MSP alert noise</a>.</p>\n\n    <h2>The mechanics that quietly break SLA reporting</h2>\n    <p>Even sensible targets fall over on the details of how the clock runs. Four mechanics to get right in your tooling:</p>\n    <ol>\n      <li><strong>Business-hours calendars.</strong> A P3 logged at 16:55 on Friday should not breach at 09:05 on Monday because the clock ran all weekend. Every SLA needs a calendar, including client-specific holidays if you serve multiple regions.</li>\n      <li><strong>Pause states.</strong> When a ticket is waiting on the client or on a vendor, the clock should pause, and the status change should be logged so the report can prove it. Without this, your worst-looking breaches are tickets where you were waiting three days for a reply.</li>\n      <li><strong>Reassignment resets.</strong> Escalating a ticket between technicians must not reset first response. It was responded to once; the client does not care about your internal routing.</li>\n      <li><strong>Breach visibility before the breach.</strong> A report that shows last month's misses is an autopsy. What changes behaviour is a queue that shows tickets at 75% of their target while there is still time to act.</li>\n    </ol>\n\n    <h2>Report on it before your client asks</h2>\n    <p>An SLA you do not measure is a marketing line, and one you only measure when challenged is a liability. Put SLA attainment on the monthly report and the quarterly review: percentage of tickets that met first response and resolution by priority, the misses, and what changed as a result. Bringing your own misses to the table, with causes, is disarming and builds far more trust than a suspiciously perfect scorecard.</p>\n    <p>Expectations around response times should also be set on day one of a new contract, alongside how tickets are raised and what counts as urgent. It is one of the steps in <a class=\"inline\" href=\"/blog/msp-client-onboarding-checklist\">our client onboarding checklist</a>, because an SLA explained in week one prevents the escalation call in month six.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>The Helios service desk was built around the mechanics above: per-client SLA policies with business-hours calendars, pause states for waiting-on-client time, and countdown timers on the queue so technicians see what is approaching breach rather than what already breached. Helio, the AI layer, auto-triages inbound tickets against your impact matrix, which keeps priorities consistent when they arrive at 2am or in capital letters. Attainment then flows into the client portal and monthly reports without anyone assembling a spreadsheet.</p>"},{"slug":"msp-client-onboarding-checklist","title":"MSP client onboarding checklist: the first 30 days done properly","meta_description":"A practical MSP client onboarding checklist: discovery, agent rollout, baselining patching, security and backups, and the 30-day review that sets margin.","published_at":"2026-07-23T00:00:00","body_html":"<p class=\"lede\">Every MSP client onboarding checklist promises a smooth first month. Most of them are really a list of forms. The estates that turn into unprofitable contracts are rarely lost at the sales stage: they are lost in the first 30 days, when unknown devices, inherited misconfigurations and unspoken expectations quietly become your problem. Here is how to onboard a new client estate so that does not happen.</p>\n\n    <h2>Onboarding is where the margin is decided</h2>\n    <p>The economics are simple. Every device you did not know about, every server nobody mentioned, every backup job that has been failing since before you arrived becomes a reactive ticket later. Reactive tickets are the most expensive kind of work an MSP does, and an estate you never properly baselined generates them for the entire life of the contract.</p>\n    <p>Industry surveys put a typical onboarding at 40 to 80 hours per client. That sounds like a lot until you compare it with the alternative: an engineer discovering, mid-incident and six months in, that the \"backed up\" file server was never in any backup job. The hours get spent either way. Onboarding is just the version where you choose when.</p>\n\n    <h2>Before day one: agree what good looks like</h2>\n    <p>The technical work goes wrong when the commercial groundwork is missing, so settle four things in writing before an agent touches a machine:</p>\n    <ul>\n      <li><strong>Scope and exclusions.</strong> Which sites, devices and applications are covered, and, just as importantly, which are not. The printer nobody owns will find you otherwise.</li>\n      <li><strong>Response targets.</strong> Realistic SLAs by priority, plus the escalation path when something breaches them. Vague promises made in a sales call become disputes in month three.</li>\n      <li><strong>Access and offboarding the old provider.</strong> Domain admin credentials, Microsoft 365 global admin, firewall logins, DNS registrar, backup consoles. Get them transferred, tested and rotated. The outgoing MSP's accounts should be disabled on a date everyone has agreed.</li>\n      <li><strong>A named contact on each side.</strong> One person who can approve changes, and one engineer who owns the onboarding end to end.</li>\n    </ul>\n\n    <h2>Week one: deploy agents and trust nothing on the spreadsheet</h2>\n    <p>The handover document says 47 devices. It is wrong. It is always wrong. Machines get bought outside procurement, old servers keep running because nobody is brave enough to switch them off, and the finance director's home laptop has a VPN profile nobody remembers creating.</p>\n    <p>So the first technical job is discovery, not configuration. Roll your monitoring agent out to everything you can reach, then reconcile what reports in against what you were told exists. The gap between those two lists is your first deliverable to the client: most of them have never seen an honest inventory of their own estate, and showing them one builds more trust in week one than any slide deck.</p>\n    <p>While the agents check in, capture the boring but vital facts per device: OS version and support status, hardware age, disk health, local admin accounts, and which machines are servers in function even if they are desktops in form factor. That last category is where future incidents live.</p>\n\n    <h2>Baseline the risk: patching, protection and backups</h2>\n    <p>With the inventory real, baseline the three things most likely to hurt you later.</p>\n    <p><strong>Patching.</strong> Record the current patch level of every device before you change anything, so you can show movement. Expect the Windows side to look better than it is and the third-party side to be worse than anyone admits: browsers, PDF readers and conferencing tools sit outside Windows Update and are usually months behind. We covered why in <a class=\"inline\" href=\"/blog/third-party-patching-for-msps\">our guide to third-party patching for MSPs</a>.</p>\n    <p><strong>Endpoint protection.</strong> Do not ask whether antivirus is \"installed\". Ask, per device: is it running, is it up to date, is real-time protection actually on, and when did it last report in? Inherited estates routinely contain machines where the AV is present but disabled, expired or pointing at a management server that no longer exists.</p>\n    <p><strong>Backups.</strong> List every backup job, confirm each one has succeeded recently, and check the coverage in both directions: jobs that are failing, and machines that matter but appear in no job at all. A green console means very little on its own, as we argued in <a class=\"inline\" href=\"/blog/backup-monitoring-for-msps\">why a green tick is not a restore</a>. If you do nothing else in week two, run one test restore.</p>\n\n    <div class=\"callout\"><p><strong>Write the baseline down and date it.</strong> A one-page snapshot of patch levels, protection gaps and backup coverage on the day you took over is the cheapest insurance an MSP can buy. It separates the problems you inherited from the problems you caused, and it turns your first quarterly review into a progress report instead of a defence.</p></div>\n\n    <h2>A practical MSP client onboarding checklist</h2>\n    <p>Condensed into one list you can lift into your PSA as a project template:</p>\n    <ol>\n      <li>Contract signed, with scope, exclusions and SLAs explicit.</li>\n      <li>All credentials transferred, tested, rotated and stored in your documentation system.</li>\n      <li>Previous provider's access disabled on an agreed date.</li>\n      <li>Monitoring agents deployed to every reachable device.</li>\n      <li>Discovered inventory reconciled against the handover list; the gap reported to the client.</li>\n      <li>Patch, endpoint protection and backup baselines captured and dated.</li>\n      <li>Critical gaps (unsupported OS, unprotected machines, unbackuped servers) raised with the client in writing, with a remediation plan.</li>\n      <li>Alerting tuned before go-live, not after.</li>\n      <li>Service desk live: users know how to raise tickets, and the client portal or email route is tested.</li>\n      <li>30-day review booked in the calendar before onboarding ends.</li>\n    </ol>\n    <p>Point eight deserves a sentence more. Switching a monitoring platform on across an unfamiliar estate with default thresholds will bury your team in noise for a fortnight, and a noisy first fortnight teaches technicians to ignore the new client's alerts. Tune thresholds against the baseline you just captured; <a class=\"inline\" href=\"/blog/msp-alert-fatigue\">our piece on alert fatigue</a> covers how to do that without missing real incidents.</p>\n\n    <h2>The 30-day review closes the loop</h2>\n    <p>Onboarding ends with a short, honest review meeting: here is what we found, here is what we fixed, here is what remains and who owns it. Bring the dated baseline and the current numbers next to each other. Clients rarely remember the individual tickets from their first month, but they remember being shown, in plain numbers, that their estate is measurably safer than when you arrived. That meeting is also where the first quarterly business review gets booked, which keeps the relationship proactive from the start.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Most of the checklist above is discovery and baselining, which is exactly the work a platform should be doing for you. Deploy the Helios agent across a new estate and the inventory, patch status, endpoint protection state and backup job health arrive in one view as devices check in, so the baseline snapshot is a report you export rather than a spreadsheet you build. Helio, the AI layer, then investigates the anomalies it finds in an unfamiliar estate rather than simply alerting on them, which matters most in exactly this period, when your team does not yet know what normal looks like for the client.</p>"},{"slug":"smes-are-not-small-enterprises","title":"SMEs are not small enterprises: what the big IT tools keep getting wrong","meta_description":"Why the big IT platforms are built for the biggest buyers, how small organisations differ from enterprise, and what that means for in-house teams and MSPs.","published_at":"2026-07-22T00:00:00","body_html":"<p class=\"lede\">I have spent my career on both sides of a line most people only see one side of. By day I have spent years in senior leadership at some of the world's largest technology and enterprise storage organisations, where downtime is measured in money per minute and an auditor can ask you to prove anything. In the evenings and weekends, I run a small MSP, where I am the one carrying the phone. Sitting in both chairs has made one thing obvious: whether you look after IT for clients or you are the in-house IT for one business, the tools you are handed were designed for an organisation that looks nothing like yours.</p>\n\n    <h2>The big platforms are built for the biggest buyer in the room</h2>\n    <p>Watch how the large IT platforms actually make decisions and it all makes sense. Their best customers are the corporate IT departments and the fifty-technician MSPs, so that is who the product serves. You see it in the pricing, which meters per technician or per admin seat and turns a growing team into a growing bill. You see it in the configuration surface, hundreds of policies and toggles that assume someone whose whole job is administering the tool. You see it in the add-on menu, where the feature you thought you were buying turns out to live behind a second line item. And you see it in the sales motion, built around demos, onboarding calls and annual commitments.</p>\n    <p>None of that is stupid. It is a rational response to where their revenue comes from. The problem is that a two-person MSP, or the one IT manager looking after a hundred-user business, inherits a platform tuned for an organisation two orders of magnitude larger, and pays for the privilege in money, time and complexity they do not have.</p>\n\n    <h2>An SME is not a small enterprise</h2>\n    <p>The tempting assumption, and the one the big tools quietly make, is that a small business is just an enterprise with the numbers turned down. It is not. The physics are different.</p>\n    <ul>\n      <li><strong>There is no tooling admin.</strong> Nobody's job is to configure the platform. The person setting up alerting is the same person answering the ticket, chasing the supplier and driving to site. Every hour of setup is an hour not spent on the work.</li>\n      <li><strong>There is no procurement buffer.</strong> The person choosing the tool is the person who has to justify it, out of a margin or a budget small enough to feel. A surprise add-on is not a line in someone else's spreadsheet, it is money the business notices this month.</li>\n      <li><strong>Downtime is personal.</strong> When a server falls over, there is no follow-the-sun team. It is you, at 2am, with somebody's business on your shoulders. The tool needs to help you fix it, not give you a dashboard to admire it in.</li>\n      <li><strong>Every hat sits on one head.</strong> Monitoring, tickets, security, backups, Microsoft 365, remote access, reporting. An enterprise has a team per box. A small organisation has one person for all of them, so the tool that stitches them together is worth more than any single feature.</li>\n    </ul>\n    <p>Give that person an enterprise platform and you have not empowered them, you have handed them a second job. What they need is not fewer capabilities. It is the same capabilities without the assumption that someone else will operate them.</p>\n\n    <h2>What enterprise actually taught me, and what it did not</h2>\n    <p>Here is the part people get wrong when they hear \"built by someone from the enterprise world\". They assume it means enterprise bloat, more process, more configuration, more meetings. It means the opposite.</p>\n    <p>What genuinely transfers from running technology at serious scale is the discipline. When downtime is not an option and an auditor can ask you to prove your controls, you learn to build things that do not fall over quietly, that fail loudly and safely, that keep one customer's data provably separate from the next. You learn that reliability is a feature you design in, not a thing you hope for. That standard is worth bringing to a small team, whether it is an MSP or an in-house one.</p>\n    <p>What does not transfer is the weight. The committees, the change-advisory theatre, the twelve-week rollout for a config change. The trick is to keep the standard and throw away the ceremony, and you can only do that if you have actually lived on both sides and know which is which.</p>\n\n    <div class=\"callout\"><p><strong>The test I hold our own product to:</strong> would this survive an enterprise audit, and could a one-person IT team switch it on before lunch? If the answer to either is no, it is not finished.</p></div>\n\n    <h2>And then there is the vibe-coding wave</h2>\n    <p>There is a lot of noise right now about how anyone can build software with AI, and to be fair, a lot of it is true. You can vibe-code a slick dashboard and a clever chat assistant in a weekend. The AI layer, the part everyone is excited about, is genuinely the easy bit now.</p>\n    <p>But a platform you run a whole IT estate on is not a dashboard. The hard part was never the AI or the pretty screens. It is the plumbing nobody demos: an agent that actually runs on every Windows, Mac and Linux machine, survives reboots and updates, patches, and lets you take control when a user is stuck. Deep, careful integrations into Microsoft 365, Defender and your backup vendors. Multi-tenant security that holds up when it matters. The reliability to still be right at 2am when something is genuinely on fire. That is months and years of unglamorous work, and no amount of AI writes it for you, because most of it is judgement about what actually matters, learned by doing the job.</p>\n    <p>So I am relaxed about the wave. Point tools that solve one problem with AI will keep appearing, and some will be good. But the thing you can hand your whole business to is not a weekend build, and it never was. If anything, the flood of shallow tools makes depth more valuable, not less.</p>\n\n    <h2>What I actually want from a tool</h2>\n    <p>When I sit in the chair of the person who actually has to run it, the list is short. One price I can plan around, with no per-technician tax and no surprise renewal letter. Every feature included, because I do not have time to work out which tier hides the thing I need. An AI that does the boring triage so my evenings are my own. And underneath it all, the quiet confidence that it was built to a standard, by people who have actually had to fix the thing at 2am.</p>\n    <p>I could not find that tool, so I built it. That is really the whole story behind Helios. Not a startup guessing from a whiteboard at what small IT teams and MSPs want, but the platform I wanted to pay for, built by someone who runs on it every day and holds it to the standards the enterprise world taught me were not optional.</p>"},{"slug":"backup-monitoring-for-msps","title":"Backup monitoring: why a green tick is not a restore","meta_description":"A successful backup job is not proof you can restore. A practical guide for MSPs and in-house IT: silent failure modes, the 3-2-1-1-0 rule and restore tests that fit a real week.","published_at":"2026-07-22T00:00:00","body_html":"<p class=\"lede\">Every IT team has a story about the backup that was green every night for six months and then would not restore. Whether you run an MSP or look after one estate in house, backup monitoring is usually treated as a solved problem: the job ran, the tick is green, move on. But \"the job ran\" and \"the data is recoverable\" are two different claims, and the distance between them is where businesses get hurt.</p>\n\n    <h2>The claim a green tick actually makes</h2>\n    <p>When Veeam, Acronis or any other backup product reports success, it is telling you one thing: the job completed without a fatal error. That is worth knowing, but notice everything it does not say. It does not say the backup covered the data that matters. It does not say the backup chain is intact or that the storage it landed on is healthy. It does not say the copy is free of the ransomware that has been sitting quietly in the environment for a fortnight. And it says nothing at all about whether anyone could actually restore it under pressure.</p>\n    <p>The failure modes that end up in incident reports are rarely loud. They are the quiet ones:</p>\n    <ul>\n      <li><strong>The job that stopped running.</strong> A schedule gets disabled during maintenance, a service account password expires, and no failure alert fires, because nothing failed. Nothing happened at all.</li>\n      <li><strong>The scope that drifted.</strong> The job still protects the file server it was pointed at in 2023. The line-of-business database that moved to a new VM last year is not in any job.</li>\n      <li><strong>The chain that broke.</strong> Incrementals keep succeeding against a corrupt full. Every night is green, and every night the estate is one bad block away from an unrecoverable chain.</li>\n      <li><strong>The retention that silently shrank.</strong> The repository filled up, old restore points were pruned, and the \"30 days of history\" in the contract or the policy quietly became six.</li>\n    </ul>\n    <p>None of these show up as a red job. All of them show up on the day you need a restore.</p>\n\n    <h2>3-2-1 was never the whole answer, and it has grown two digits</h2>\n    <p>The classic 3-2-1 rule still holds: three copies of the data, on two different media, with one copy off site. It is a good test of architecture. But it describes where copies live, not whether they work, which is why the industry has largely moved to <strong>3-2-1-1-0</strong>: the extra 1 is one copy that is offline or immutable, so ransomware cannot encrypt the backups along with the production data, and the 0 is <strong>zero errors on verification</strong>, meaning backups are tested and confirmed restorable, not assumed.</p>\n    <p>That final zero is the part most backup setups are missing, in service providers and in-house teams alike. Architecture is a design decision you make once. Verification is an operational habit you keep every week, and habits are harder to sell, harder to schedule and easier to drop when the queue is full. Which is exactly why it should be systematised rather than left to good intentions.</p>\n\n    <h2>What backup monitoring should actually check</h2>\n    <p>A monitoring setup you can stand behind checks four layers, in order of how cheap they are to automate:</p>\n    <ol>\n      <li><strong>Job outcomes, including silence.</strong> Alert on failures, obviously. But also alert when an expected job simply does not report. \"No result in 26 hours\" must be treated as seriously as \"failed\", because the stopped job is the more dangerous of the two.</li>\n      <li><strong>Coverage against inventory.</strong> Reconcile the list of protected machines against the list of machines that exist. Every device in the estate should either be in a backup job or explicitly marked as not requiring one. New servers and rebuilt VMs are how scope drift starts.</li>\n      <li><strong>Restore point age and retention.</strong> Track the age of the newest restore point per protected system, and the depth of history, against what your contract or your own recovery policy promises. A green job with a three-week-old restore point is a red finding.</li>\n      <li><strong>Verification results.</strong> Use the vendor's built-in verification where it exists, health checks and scheduled verify jobs in Veeam, validation in Acronis, and record the results centrally. A verification job that never runs is a policy, not a control.</li>\n    </ol>\n\n    <h2>Restore tests that fit in a real week</h2>\n    <p>Full disaster recovery rehearsals matter, but if the bar for testing is \"rebuild an entire site from bare metal\", testing will happen once a year at best. The practical answer is a tiered cadence that trades depth for frequency:</p>\n    <ul>\n      <li><strong>Weekly, automated:</strong> vendor verification jobs and a scripted file-level restore of a handful of files from a random protected machine, checksummed against the source.</li>\n      <li><strong>Monthly, lightweight:</strong> boot one server backup as an isolated VM and confirm the operating system and key application actually start. Fifteen minutes per site, rotated so every critical system is exercised over a quarter.</li>\n      <li><strong>Twice a year, per critical client or department:</strong> a timed restore of the most important workload, with the result written down: how long it took, what was missing, what surprised you. That number is your real recovery time, whatever the recovery plan says.</li>\n    </ul>\n\n    <div class=\"callout\"><p><strong>Make the evidence do double duty.</strong> Every restore test produces proof: a screenshot of the booted VM, a checksum match, a timed run. Save it and put it in the next review, whether that is a client QBR or an internal update to the board. \"We restored the finance server in 41 minutes last month\" is the most persuasive sentence you can say about backup, and it costs nothing extra once testing is routine.</p></div>\n\n    <h2>Turn backup state into daily ops, not a monthly report</h2>\n    <p>The last piece is organisational. Backup health should be visible in the same place your technicians already work, not in a separate console per vendor that someone checks when they remember. A failed or silent job should become a ticket with an owner and an SLA, the same as a down server. The counterweight is discipline about noise: one warning per chain of related failures, not forty emails from four consoles. The techniques in our guide to <a class=\"inline\" href=\"/blog/msp-alert-fatigue\">cutting alert noise without missing real incidents</a> apply to backup alerts more than anything else, because backup noise is precisely the kind that trains people to stop reading.</p>\n    <p>Get those pieces in place, coverage reconciliation, silence detection, retention tracking, tiered restore tests and single-pane visibility, and backup stops being a leap of faith. It becomes a control with evidence behind it, which is what the business you support assumes it already has. If you are an MSP, that evidence matters twice over, because you are also the one who carries the liability when a restore fails.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>Helios monitors Veeam and Acronis backups from the same agent that already watches patching, antivirus and system health, so backup state sits in the one dashboard your team actually looks at, across every site or client you look after. Missed and silent jobs surface as alerts rather than as absences, and Helio, the AI layer, folds backup findings into device investigations, so \"when did this machine last have a good restore point\" is a question with an instant answer rather than a console safari. The restore testing habit is still yours to build, but the watching, reconciling and chasing is work software should be doing.</p>"},{"slug":"third-party-patching-for-msps","title":"Third-party patching: why it is harder than Windows Update","meta_description":"Browsers, PDF readers and conferencing tools are where most breaches start, yet they patch on their own schedules. A practical look at third-party patching at scale, winget's sharp edges, and what real coverage requires.","published_at":"2026-07-21T00:00:00","body_html":"<p class=\"lede\">Windows Update is the part of patching everyone can see. It has a schedule, a dashboard and a reboot prompt. The software that actually gets exploited, browsers, PDF readers, conferencing tools, updates on its own terms, in the background, one machine at a time. Covering that at scale is where most patch strategies quietly come apart, whether the machines belong to clients or to your own company.</p>\n\n    <h2>The gap nobody reports on</h2>\n    <p>Look at where real intrusions start and you keep landing on the same short list: a browser a few versions behind, an unpatched PDF reader, an old build of a remote-access or conferencing tool. None of those are Windows. All of them sit outside Windows Update, and none of them show up on the patch report you are most likely to be watching.</p>\n    <p>That is the uncomfortable part. A fleet can read as <strong>98% patched</strong> on the Windows side and still be carrying the exact software a phishing payload is built to exploit. The number looks healthy because it is measuring the easy 20% of the problem.</p>\n\n    <h2>Why third-party is structurally harder</h2>\n    <p>Windows Update is one channel, run by one vendor, on a predictable cadence. Third-party software is the opposite of all three of those things:</p>\n    <ul>\n      <li><strong>No single channel.</strong> Every vendor ships updates their own way, on their own schedule, some silently and some not at all until a human clicks.</li>\n      <li><strong>Per-user sprawl.</strong> A lot of apps install per user, not per machine. The same laptop can carry three copies of the same tool for three profiles, each a different version.</li>\n      <li><strong>Version drift.</strong> Nobody has a clean inventory of what is installed across the estate they look after, let alone what the current safe version of each thing is.</li>\n    </ul>\n    <p>This is why \"we patch monthly\" almost always means \"we patch Windows monthly.\" The third-party side needs a different mechanism, and for a while there was not a good one that worked without an agent per application.</p>\n\n    <h2>winget helped, but it has sharp edges</h2>\n    <p>The Windows Package Manager, <code>winget</code>, changed the picture. It can enumerate installed apps that have an update available and install them silently, from a curated source, without a per-app tool:</p>\n    <pre><code>winget upgrade --source winget --accept-source-agreements\nwinget install --id Google.Chrome --exact --silent \\\n  --accept-package-agreements --accept-source-agreements</code></pre>\n    <p>For a technician at a keyboard, that is close to magic. Run it unattended across hundreds of machines, whether that is an MSP covering a dozen client tenants or one internal team covering every laptop in the building, and it comes with edges worth knowing before you trust it:</p>\n    <ul>\n      <li><strong>Execution context matters more than you think.</strong> Monitoring agents run as a background service under the LocalSystem account. The friendly <code>winget</code> on a user's PATH is a per-user alias that does not exist for LocalSystem, so a naive call simply fails to find it.</li>\n      <li><strong>Per-user installs can be invisible.</strong> Run winget as the system account and it sees machine-wide installs cleanly, but software your users installed just for themselves may not appear at all.</li>\n      <li><strong>It will wait for a human.</strong> If any step wants confirmation, an unattended install can sit and block until something times out and looks like a failure, when nothing was really wrong.</li>\n      <li><strong>Reboots and exit codes lie.</strong> An install that succeeds but needs a restart returns a non-zero code. Treat every non-zero result as a failure and you will bury real successes in false alarms.</li>\n    </ul>\n\n    <div class=\"callout\"><p><strong>The one that bites everyone:</strong> a tool that shells out to <code>winget</code> from a system service, catches the \"not found\" error and quietly returns an empty list, will report \"no third-party updates\" on every device forever. It looks like a clean estate. It is a silent blind spot. Always be suspicious of a patching feature that has never once found anything to do.</p></div>\n\n    <h2>What real coverage requires</h2>\n    <p>Getting third-party patching right is less about the install command and more about everything around it. A dependable setup:</p>\n    <ol>\n      <li><strong>Runs in the right context</strong>, resolving the real tool path rather than assuming a user's environment.</li>\n      <li><strong>Never trusts a silent empty result.</strong> \"Found nothing\" and \"could not look\" must be different outcomes, and the second one should be loud.</li>\n      <li><strong>Handles reboots and benign exit codes</strong> as the successes they are, not failures.</li>\n      <li><strong>Reports per package</strong>, so a batch where one app fails does not hide behind a green tick.</li>\n      <li><strong>Keeps an inventory</strong> of what is installed against what is current, so the gap is visible before it is exploited, not after.</li>\n    </ol>\n    <p>Do those five things and third-party patching stops being a box that is ticked and starts being a control you can actually stand behind in a security review.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>We built third-party patching into the Helios agent for exactly the reasons above: the browsers and readers are the real attack surface, and covering them should not need a separate product per application. It runs from the same lightweight agent that already reports inventory, patches, antivirus state and backups, so the third-party view sits next to Windows Update rather than in a tool of its own. And yes, we learned the LocalSystem lesson the hard way, which is partly why this article exists.</p>"},{"slug":"msp-alert-fatigue","title":"Alert fatigue in IT: how to cut alert noise without missing real incidents","meta_description":"Alert fatigue is a configuration problem, not a staffing problem. A practical guide for MSPs and in-house IT teams on cutting alert noise at the source.","published_at":"2026-07-21T00:00:00","body_html":"<p class=\"lede\">Alert fatigue rarely arrives as a crisis, and it does not care whether you run a managed service provider or the IT team inside a single business. It creeps in one ignored notification at a time, until the day a technician swipes away the alert that actually mattered. The fix is not more discipline or more staff. Alert noise is a configuration problem, and configuration problems can be engineered away.</p>\n\n    <h2>The real cost of a noisy queue</h2>\n    <p>Every alert that lands in front of a technician makes a small claim on their attention. When most of those claims turn out to be nothing, disk usage crossing 80% on a server with 400&nbsp;GB free, a heartbeat blip during a scheduled reboot, the same offline printer for the third week running, the team learns a dangerous lesson: alerts are usually safe to ignore.</p>\n    <p>That lesson has a price. Response times stretch because nobody trusts the queue. Escalations get missed because the genuine failure looks identical to the noise around it. And your most experienced engineers spend their sharpest hours triaging notifications a rule could have closed, while the people you support wait. If you have ever found a real outage buried under forty routine alerts, you have paid it.</p>\n\n    <h2>Where RMM alert noise actually comes from</h2>\n    <p>Most alert fatigue traces back to three decisions made early and never revisited:</p>\n    <ul>\n      <li><strong>Default monitoring policies.</strong> RMM platforms ship with cautious defaults that alert on everything, because a missed condition looks worse for the vendor than a noisy one. Left unchanged, defaults generate the bulk of your queue.</li>\n      <li><strong>Instant thresholds.</strong> A CPU spike that lasts thirty seconds is normal behaviour. A CPU pegged for thirty minutes is a problem. Alerts that fire the moment a line is crossed cannot tell the difference.</li>\n      <li><strong>Snowflake configurations.</strong> When every site or client estate carries its own hand-tuned alerting, nobody can reason about the whole. A sensible exception in one place becomes a blind spot in another, and tuning work never compounds. For an MSP the effect multiplies, because one untuned default fires across every client estate at once.</li>\n    </ul>\n    <p>None of these are staffing problems. All of them are fixable in an afternoon per policy, which is why the highest-leverage monitoring work available to any IT team is boring: sit down and re-decide what deserves a human.</p>\n\n    <h2>Cut the noise at the source: five changes that work</h2>\n    <ol>\n      <li><strong>Alert on symptoms, not causes.</strong> Your users feel \"the application is down\", not \"a service stopped\". Monitor the outcome where you can, and let cause-level signals feed a ticket's context rather than raising their own.</li>\n      <li><strong>Require persistence.</strong> Add a duration to every threshold: disk above 90% for 24 hours, host offline for 10 minutes, service down after one failed auto-restart. Transient conditions should resolve themselves silently.</li>\n      <li><strong>Standardise your baselines.</strong> Run one monitoring baseline per device role, applied everywhere: across your clients, or across your own sites if you are in-house, with the exceptions written down. Tuning done once then improves the whole estate you look after.</li>\n      <li><strong>Give every alert an action.</strong> If the correct response to an alert is \"acknowledge and move on\", the alert should not exist. Either attach a runbook step, automate the response, or delete the rule.</li>\n      <li><strong>Let automation take the first swing.</strong> Restarting a stopped service, clearing a temp directory, retrying a failed backup job: if a script fixes it nine times out of ten, a human should only hear about the tenth.</li>\n    </ol>\n\n    <div class=\"callout\"><p><strong>The 90-day test:</strong> for each alert rule, ask when a technician last took a real action because of it. If the honest answer is \"not in the last 90 days\", the rule is noise wearing a safety costume. Delete it or demote it to a report.</p></div>\n\n    <h2>Triage what remains</h2>\n    <p>Even a well-tuned estate produces alerts, so the survivors need a shape. Three tiers are enough for most teams: <strong>wake someone up</strong> (an outage the people you support can feel, or data at risk), <strong>work today</strong> (degraded but standing), and <strong>review weekly</strong> (trends and hygiene). Anything that cannot be placed in a tier goes back through the 90-day test.</p>\n    <p>Then remove the duplicates. One dying switch should raise one incident, not thirty device-offline alerts. Deduplication, maintenance windows that actually suppress alerts during patching, and grouping related signals into a single ticket will do more for a technician's morning than any dashboard redesign.</p>\n\n    <h2>Measure it or it will grow back</h2>\n    <p>Alert noise regrows because every new site or client, tool and integration adds rules and nobody removes them. Track three numbers monthly:</p>\n    <ul>\n      <li><strong>Alerts per technician per day.</strong> The raw load. Watch the trend, not the absolute number.</li>\n      <li><strong>Actionable ratio.</strong> The share of alerts that led to a real action. Below roughly half, trust in the queue starts to die.</li>\n      <li><strong>Time to acknowledge for the top tier.</strong> The number that tells you whether the alerts that matter are still being believed.</li>\n    </ul>\n    <p>Put them in the same monthly review as your service targets. If you are an MSP those targets are contracted response obligations, so a queue nobody believes in is a commercial risk as well as an operational one. Either way the rule holds: a queue that is measured gets pruned, and a queue that is not measured becomes wallpaper.</p>\n\n    <h2>Where this fits with Helios</h2>\n    <p>We designed Helios monitoring around the ideas above rather than bolting them on. Baselines are standardised across every site and client by default, thresholds carry persistence out of the box, and Helio, the AI layer in the platform, auto-triages what does fire: it groups related signals, investigates the device, and auto-heals the failures that follow a known pattern, so technicians see a short queue of things genuinely worth their time. The goal is not zero alerts. It is a queue your team still believes in at 4pm on a Friday.</p>"}]}