Download PDF Full pipeline overview →

Machine Data Insights

CIM Normalization Pipeline

Turn raw machine data into validated, deploy-ready content for Splunk® ES & ITSI - faster, with AI-accelerated tooling and repeatable methods.

There’s Gold In That Data!®

The Foundation

What is CIM Normalization?

CIM normalization aligns security and operational data from every vendor to Splunk’s standard schemas - so one detection, dashboard, or KPI works across all of them, instead of a custom search per technology.

Map vendor-specific fields to common CIM fields.
Create eventtypes that identify significant security and operational events.
Create tags that link those events to the accelerated data models Splunk ES & ITSI run on.
Make every source analyzable consistently - firewall, endpoint, identity, cloud, email, PLC, SCADA, IoT.

Why It Matters

What CIM Normalization Unlocks in Splunk ES & ITSI

You can write detections without CIM - most teams have, one search per technology. Normalized data is what lets a single detection, KPI, or dashboard span every vendor, run at data-model speed, and use the content Splunk ships. ITSI KPIs and entity rules want the same thing: consistent fields and identifiers. Miss it and nothing errors - you just pay for coverage one search at a time.

For Splunk ES

One correlation search covers every vendor in a data model - not one per technology.
Out-of-the-box Security Content works as shipped.
Better detection accuracy, fewer false positives, faster investigations.
Asset & identity enrichment keys on CIM fields - it works even where detections don’t use data models.

For Splunk ITSI

One KPI definition measures every vendor's tier the same way.
Consistent identifiers mean entity rules match - hosts don't drop out of services.
Content packs depend on well-built add-ons; the pipeline produces them.
Service health scores you can actually trust - and explain to the business.

Sound familiar?

Not every sourcetype gets normalized - and we know it
Splunkbase TAs that miss sourcetypes
No inventory of what needs normalizing
Sources missing from your detections
Splunk’s own content that never fires

What breaks when you move a detection onto a data model without CIM normalization

unmappedpan:traffic
no tag
data modelNetwork Traffic

Converting a search to run on a data model is where unmapped sourcetypes surface. The data is in Splunk. The search runs clean. It finds nothing - because those events never reached the model it queries. MDI’s CIM Assessment Toolkit (CAT) finds them first, ranked by impact.

The MDI CIM Normalization Pipeline

You run the client-side tools; the encrypted exchange hands data to MDI; MDI runs the engine.

01
Assess
CAT
02
Sanitize
Paydirt
03
Secure Handoff
Secure Exchange
04
Build
Data Refinery
01 · Assess

CAT

CIM Assessment Toolkit

Free, open-source Splunkbase app that scores CIM compliance by dataset and sourcetype - exposing the gaps that break Splunk ES correlation searches.

Free & Open SourceDataset-LevelAcceleration Health
Paydirt
02 · Sanitize

Paydirt

Log Scrubber

Free, open-source scrubber for CUI, PII, PHI & credentials. Runs fully on your machine - no install, no network calls. CMMC, HIPAA & GDPR aware.

Free & Open SourceRuns OfflineNo Install
03 · Secure Handoff

Secure Exchange

Encrypted Data Transfer

A private, encrypted cloud channel per engagement - clients upload source data, collect signed deliverables. Isolated, encrypted, fully logged.

EncryptedPer-Client IsolationAudit Trail
04 · Build

Data Refinery

Normalization Engine

The MDI engine that builds the fix an assessment finds: validated, CIM-compliant Splunk TAs and Cribl packs - AppInspect-checked and deploy-ready.

CIM-ValidatedAppInspect-CheckedTAs + Cribl Packs

One Validated Output, Two Deployment Paths

Data Refinery’s CIM-validated artifacts deploy whichever way your environment runs.

Data Refinery
CIM-validated output
Data Refinery refining raw machine data into deployable artifacts

Splunk CIM TAs

props, transforms, eventtypes & tags - normalize vendor fields to CIM at search / index time. Drop-in for Splunk ES.

Cribl Packs

normalize and reduce in-stream - cut ingest and license cost before data ever lands.

Both paths ship CIM-validated, AppInspect-checked, and fully documented.

Let’s turn your data into detection.

MDI delivers CIM normalization, data-volume reduction, and CIM macro & correlation-search optimization - with AI-accelerated tooling that cuts time-to-value and consulting cost.

CIM Normalization Data-Volume Reduction Data Model Acceleration Correlation-Search Optimization CIM Macro Optimization AI-Accelerated

Scan to explore