Case Study

Azure
AI Observability
Cost Tracing

Enabling End-to-End Traceability for an Azure AI-Powered Application Using OpenTelemetry

With the rapid growth of AI applications, having clear visibility into how everything works behind the scenes has become more important than ever. In this case study, we share how we partnered with a global asset management firm to bring full end-to-end traceability to their AI application running on Azure. By combining the power of OpenTelemetry for distributed tracing with Azure Application Insights for centralized observability, we helped them gain the insights they needed to keep their systems running smoothly and reliably.

Reduction in Fail Rate
0 %
Reduction in Mean Time to Recovery
0 %
Incident Rate Reduced to
Up to 0 %

Client Profile

INDUSTRY

Global alternative asset management and retirement services, with a strong focus on credit, equity, real assets, and long‑term retirement income solutions

SCALE

Large, multinational financial organization with a diversified portfolio, long‑duration capital strategies, and a global presence

PLATFORM

Microsoft Azure

CHALLENGE

Limited end‑to‑end visibility across a complex AI application stack made it difficult to trace user prompts from entry to final response, diagnose issues, understand model behavior, and satisfy regulatory and internal reporting requirements

ENGAGEMENT

Design and implementation of a full observability and tracing architecture using OpenTelemetry and Azure Application Insights to deliver complete prompt‑to‑response traceability, faster incident resolution, and better AI cost and performance management.

Executive summary

Our client is a leading global asset management firm known for its innovative approach to alternative investments and retirement solutions. Guided by values like pushing boundaries, creating opportunities and leading with integrity, they support both institutional and individual investors across credit, equity, real assets and retirement strategies.

As part of their digital transformation journey, the firm built an AI-powered Application on Azure to enhance customer engagement—showcasing their commitment to responsible, scalable technology and modern client experiences.

Business Challenges

The client faced four critical operational and security challenges that were inhibiting growth and increasing risk:

1

Lack of End‑to‑End Visibility

There was no unified way to see the entire request lifecycle—from user prompt through search, model inference, orchestration, and final response. Teams struggled to map user journeys, identify bottlenecks, and confidently explain system behavior to stakeholders and regulators.

2

Complex Multi‑Service Architecture

The AI experience relied on tightly integrated Azure services (OpenAI, AI Search, Doc Intelligence, backend APIs, orchestration components, AKS, APIM), each generating its own telemetry and logs. Without a standard tracing model, correlating events across services was difficult, slowing root‑cause analysis and making performance optimization guesswork.

3

Operational Reliability and Compliance Risks

Failures were hard to diagnose, alerts were limited, and visibility into AI model performance and accuracy was incomplete. This led to longer mean time to detect and resolve incidents, potential user impact, and challenges in demonstrating compliance and control in a regulated financial environment.

Solution Architecture

Techanek designed a three-phase automated infrastructure deployment architecture built on native AWS services

OpenTelemetry‑Based Trace Model

A parent‑child span hierarchy was implemented for every user request, with a root span representing the initial prompt and child spans capturing each downstream step (AI Search, Doc Intelligence, OpenAI calls, backend APIs, orchestration, response assembly). Each span was enriched with metadata such as latency, function name, document details, model configuration, and correlation IDs to enable deep analysis of system behavior.

Seamless Azure Service Integration

OpenTelemetry SDKs and Azure‑compatible exporters were used to send traces, metrics, and logs into Azure Monitor and Application Insights. This created centralized dashboards and visualizations where teams could see full request flows, performance hotspots, and error patterns across all Azure components.

Enhanced Debugging and Cost/Token Visibility

Telemetry was extended with error flags, exception details, token usage data, execution metrics, and business context (e.g., which documents were retrieved, which models were invoked). This enabled faster root‑cause analysis and allowed attribution of AI model costs to specific workloads and usage patterns, supporting better forecasting and FinOps for AI services.

Integration with Existing AI Application Flow

Tracing was embedded directly into the AI request pipeline: prompt ingestion, context retrieval via AI Search, document parsing via Doc Intelligence, model invocation, response generation, and persistence of trace data. Traces were correlated end‑to‑end from the user interaction in the AI application front‑end through the gateway, AKS cluster, and Azure services back to Application Insights.

Technology stack

Cloud Services

Azure OpenAI
Azure AI Search
Azure Doc Intelligence
Azure Blob Storage

Observability and Tracing

OpenTelemetry
Azure Monitor
Azure Application Insights
Azure Workbooks

Business Outcomes

The automation transformation delivered measurable impact across security, efficiency, and scalability dimensions:

Operational Reliability
BEFORE

Incidents in the AI application were difficult to diagnose, with limited visibility across services leading to longer outages and inconsistent user experience.

AFTER

End‑to‑end tracing reduced fail rates by 60% and cut mean time to recovery by 75%, giving teams faster, clearer paths to resolution and more stable operations.

Change Management and Delivery Speed
BEFORE

Understanding the impact of changes to models, search configuration, or backend workflows required manual investigation and carried higher risk.

AFTER

The span‑based trace model reduced lead time for changes by 85%, enabling safer experimentation and faster iteration with trace‑back to affected components.

Incident Frequency and User Experience
BEFORE

Limited proactive monitoring and AI‑specific observability resulted in recurring issues that affected users before teams had enough context to respond.

 

AFTER

Better alerts, richer telemetry, and consolidated dashboards led to a 40% reduction in incident rate and a more consistent, trustworthy AI experience.

Cost and Token Usage Governance
BEFORE

AI model token usage and observability spending were hard to track and optimize across services.

AFTER

OpenTelemetry‑driven telemetry allowed correlation of token usage and model calls with specific traces, contributing to around 30% lower observability costs and improved AI cost management.

Key Capabilities Delivered

Every user request is now represented as a structured trace with parent/child spans, covering search, model inference, orchestration, and response delivery.

Centralized telemetry in Application Insights provides a single pane of glass for metrics, logs, and traces across the entire AI stack.

Rich span metadata and correlation IDs make it significantly easier to pinpoint failing components and performance bottlenecks.

Telemetry captures model usage and performance in context, supporting better resource planning, AI governance, and continuous optimization

Conclusion

Implementing end-to-end tracing with OpenTelemetry transformed how our client understood and managed their Azure AI Application. By capturing every step from prompt to response they gained clear visibility into the system, making it easier to fix issues, improve performance, and deliver more accurate answers. With trace data now flowing into Azure Application Insights, the team can quickly respond to feedback, ensure reliability and keep improving turning observability into a key driver of trust and continuous improvement.

Ready to make your AI systems fully transparent and traceable?