메뉴 바로가기 본문 바로가기

티맥스소프트

전체 검색 입력 폼

Blog / News

  1. HOME
  2. About
  3. Blog / News
[Op-Ed] Beyond AIOps: Enterprise IT Operations Enter the Era of AI Resilience
AI is fundamentally different from conventional software. Traditional software was designed and valued for its ability to operate reliably according to predefined specifications. The question was largely one of functionality: how many features it had and how consistently they worked. IT operations, in turn, depended on operators following predetermined procedures and methods to keep systems in a normal state.

 

The Odyssey, one of humanity's great classics, is being reinterpreted once again on the back of a recent box-office hit. Odysseus is often remembered as the strategist who led the Greeks to victory in the Trojan War. Yet the true arc of his story begins only after the war ends, with his long journey home.

 

He won a war decisive enough to bring down an entire civilization, using the Trojan horse tactic that remains a metaphor for IT security attacks today. And still, he could not return home. He lost his ships. He lost his companions. At one point, he even lost his memory, crossing the threshold of death more than once. Had the story ended as a simple tale of military victory and flawless strategy, it would not have continued to captivate readers 3,000 years later.

 

Odysseus's journey is a continuous cycle of adapting, recovering, and setting out again amid crises that never stop arising. Perhaps that is why the story so readily calls to mind the importance of resilience in enterprise IT environments. Human history itself has often followed the same pattern.

 

For decades, the ideal of enterprise IT was clear: build systems that do not fail and that perform at a high level. Organizations invested in better hardware, more stable software, and increasingly sophisticated operating systems for one overriding reason: to reduce IT failures and sustain business continuity.

 

Over time, this concept of failure management expanded into IT resilience. If failures can never be prevented entirely, then organizations need the capabilities and mechanisms to respond and recover quickly. This does not diminish the importance of preventing failures. Rather, it underscores how critical rapid restoration becomes once a failure occurs.

 

With the arrival of the artificial intelligence era, the scope of IT resilience is expanding further. Advances in generative AI and AI agent technologies are accelerating the adoption of AIOps-based operational automation. In an AIOps environment, AI agents embedded in enterprise operations management tools can install software, configure environments and settings, analyze failures and events, and recommend remediation steps. Enterprises expect these capabilities to deliver higher productivity, faster decision-making, greater operational efficiency, and lower operating costs.

 

Reality, however, is more complicated than it appears, especially as IT resilience extends into AI resilience.

 

AI is fundamentally different from conventional software. Traditional software was designed and valued for its ability to operate reliably according to predefined specifications. The question was largely one of functionality: how many features it had and how consistently they worked. IT operations, in turn, depended on operators following predetermined procedures and methods to keep systems in a normal state.

 

The large language models (LLMs) on which AI agents rely do not always behave in such predetermined ways. AI agents can make judgments and take actions that differ from those of human operators, while processing large volumes of work at speeds far beyond human capacity.

 

Work performed by AI agents is shaped by the workflows coded into the agents themselves, as well as by the prompts and contextual information given to the AI model. Yet the AI model that serves as the agent's brain sits on an ambiguous boundary: it appears to think and judge for itself, while also probabilistically combining data that resembles what it has learned. It is on this boundary that AI agents have evolved from entities that answer into entities that act. As a result, the probabilistic problem commonly known as hallucination is always present.

 

The real issue is not that AI can make mistakes. It is whether we can recognize those mistakes when they occur.

 

That is why IT operations built on AIOps in the age of agentic AI must consider not only how trustworthy an AI's operational actions are, but also how those actions can be controlled and monitored so that irreversible consequences do not occur. This is where AI governance and AI guardrails become essential.

 

Traditional IT resilience focused on how quickly systems could be restored after a failure. AI resilience goes further. It asks how quickly an AI error can be detected, how the impact of an AI's incorrect action can be minimized, and how AI itself can be used to defend against AI-driven attacks. In other words, AIOps capabilities that analyze operational data, detect anomalies, and identify root causes must be designed not only for automation, but also from the perspective of operational resilience.

 

If today's AIOps is aimed at detecting failures faster and improving operational efficiency, tomorrow's AIOps must evolve into a platform for designing resilience. It must also secure visibility across AI models, agents, applications, and data pipelines as a whole.

 

Enterprises need systems that define approval stages for AIOps agent execution authority and that can perform root-cause analysis and appropriate responses when anomalous behavior occurs. Observability that can explain AI actions, policy-based governance, and security must work together as part of the same operational foundation.

 

Odysseus moved forward not by controlling every crisis he encountered, but by overcoming them one by one. Enterprises in the AI era face a similar reality. They cannot eliminate every IT failure, security breach, or AI error in advance. What they can do is design for resilience. In the AI era, innovation without trust and productivity without stability both lose their meaning.

 

Odysseus remains compelling not because he eliminated every danger, but because he overcame the countless crises of his ten-year journey home. The same is true for enterprises operating in an AI era where IT failures, security breaches, and AI errors cannot be completely prevented. The enterprises that stand out will be those that adapt to the new paradigm of IT operations: the shift beyond AIOps toward AI resilience.
 

*This article is an English translation of an op-ed by KiEun Park, CTO (EVP) of TmaxSoft, originally published in Electronic Times on September 8th, 2026.

Original op-ed: [기고] 기업 IT 운영, AIOps 넘어 'AI 복원력' 시대 온다