Skip to content
  • Sanjay 

Voice Security: The Enterprise Attack Surface Nobody Secured

Voice Security or the Phone has become one of the most overlooked attack surfaces in enterprise cybersecurity.

For years, organizations invested heavily in securing email, endpoints, identities, applications and cloud infrastructure.

The phone was different.

It was considered an operational system. A communications channel. Something the telecom team managed and the security team occasionally monitored.

That assumption is becoming expensive.

Vishing activity increased dramatically in recent years. Assertion’s Voice Threat Analysis 2026 reports a 442% rise in vishing attacks between the first and second half of 2024, estimates $44.5 billion in contact-center fraud exposure in 2025, and reports that approximately 1 in 127 calls reaching a contact center are fraudulent. The report also highlights the growing impact of deepfake attacks, with 62% of organizations surveyed reportedly experiencing one within a 12-month period.

The numbers tell only part of the story.

The bigger problem is architectural:

Most enterprises have security controls around the systems their employees use. Far fewer have security controls around the calls those employees receive.

That makes voice the next major security perimeter.


What Is Voice Security?

Voice security is the practice of protecting enterprise phone systems, contact centers, SIP infrastructure and voice applications from malicious, fraudulent, abusive and unwanted voice activity.

It includes protection against threats such as:

  • Vishing and voice phishing
  • Social engineering
  • Robocalls
  • Scam calls
  • Toll fraud
  • Telephony Denial of Service (TDoS)
  • Caller impersonation
  • IVR attacks
  • Voice-based reconnaissance
  • AI-generated voice attacks
  • Deepfake and voice-cloning attacks
  • Malicious or abusive calling patterns

Voice security is different from traditional telecom monitoring.

Monitoring tells you what happened.

Voice security is about determining whether a call should be trusted and what should happen next.

That distinction is becoming critical.


Why Voice Has Become a Cybersecurity Attack Surface

Ask a CISO what their organization secured over the last few years.

The answer will probably include:

Email.
Identity.
Endpoints.
Cloud.
Applications.
Data.

Now ask:

“What happens when a malicious caller reaches your contact center?”

The answer is often much less obvious.

That is the blind spot.

An attacker does not necessarily need to compromise an endpoint.

They can call an employee.

They can call an agent.

They can call an IVR.

They can call a patient room.

They can call a customer-service line.

They can call hundreds of extensions to discover which ones exist.

They can keep a line occupied.

They can attempt to manipulate an agent into revealing information.

They can impersonate a customer, executive, vendor or regulator.

And increasingly, they can use AI to make the interaction more convincing.

The phone has therefore evolved from a communications channel into an initial-access and fraud channel.

The cybersecurity community already recognizes voice as an attack vector. CISA’s ATT&CK-aligned material categorizes Spearphishing via Voice as an Initial Access technique.

The enterprise security perimeter is expanding.

And the telephone is part of it.


The Human Is the New Attack Surface

Traditional cybersecurity often focuses on protecting systems from malicious code.

Voice attacks frequently target something harder to firewall:

human trust.

A convincing caller can create urgency.

They can establish authority.

They can create fear.

They can sound familiar.

They can claim to know information about the person they are calling.

They can exploit the fact that a human being is naturally inclined to believe a conversation that sounds legitimate.

This is why vishing is so dangerous.

The attacker is not necessarily trying to break the phone system.

They are trying to make the person on the other end do something.

And AI is making that easier.


AI Has Changed the Voice Threat Landscape

Voice cloning and generative AI have lowered the barrier to convincing impersonation.

An attacker can potentially combine:

  • Publicly available voice recordings
  • Social media information
  • Organizational information
  • AI-generated scripts
  • Automated calling
  • Caller-ID manipulation
  • Human social-engineering techniques

The result is a more targeted attack.

Assertion’s Voice Threat Analysis 2026 describes the shift as a move toward attacks that are fewer, sharper and better researched, rather than simply larger volumes of indiscriminate fraud.

That matters.

Because a security strategy built only around blocking large volumes of obvious spam may miss the most dangerous calls.

The next-generation threat may be:

one call, to one person, at exactly the right time.


Voice Authentication Is Not the Same as Voice Trust

One of the most important distinctions in enterprise voice security is the difference between authentication and trust.

Technologies such as STIR/SHAKEN help improve caller-ID authentication.

But authentication answers a relatively narrow question:

“Is this telephone number associated with the originating caller in the way the network says it is?”

Security needs to answer a broader question:

“Should this interaction be trusted?”

Those are not the same thing.

The FCC has described STIR/SHAKEN authentication information as one component of broader call analytics and risk assessment used by providers to identify potentially unwanted or fraudulent calls.

That distinction is critical.

A call can have legitimate signaling and still involve:

  • Social engineering
  • Account takeover
  • Fraud
  • Impersonation
  • Insider manipulation
  • Abuse of an IVR
  • Reconnaissance

A trusted telephone number does not automatically mean a trusted conversation.


Why SBCs Alone Are Not Enough

The Session Border Controller remains a critical component of enterprise voice architecture.

But an SBC is primarily responsible for controlling and securing voice sessions at the network boundary.

Modern voice security requires an additional layer of intelligence.

The question is no longer simply:

“Can this call enter my network?”

It becomes:

“What do I know about this call, how risky is it, and what should happen to it?”

This is where an intelligent voice security layer becomes valuable.

Assertion has long positioned this distinction in its VoIP security research: voice infrastructure has security characteristics that cannot simply be addressed by conventional IT security tools, because voice protocols, signaling and calling behavior require specialized controls.

The architecture therefore becomes:

Carrier → SBC → Voice Security Layer → Contact Center / UC / IVR

05_modern_voice_security_architecture

The security layer can inspect calls, apply policies, analyze behavior and determine the appropriate response.


What Does a Voice Security Platform Actually Do?

A modern enterprise voice security platform should be able to perform five fundamental functions.

1. Detect

Identify suspicious calls and calling behavior in real time.

This can include:

  • Known malicious numbers
  • Abnormal calling patterns
  • High-volume calling
  • Repeated short calls
  • Unusual destination behavior
  • Suspicious signaling
  • Known threat intelligence
  • First-seen callers
  • Pattern anomalies

2. Analyze

No single signal is sufficient.

A voice security platform should combine multiple signals to establish a risk profile.

This is where AI and behavioral analytics become important.

Instead of:

Known bad = block

the model becomes:

Multiple suspicious signals = increasing risk

3. Decide

Once the call has been assigned a risk level, the system should determine what happens next.

For example:

Low risk → Allow

Medium risk → Challenge

High risk → Route / isolate / alert

Known malicious → Block

4. Respond

The response should happen in real time.

That could mean:

  • Blocking the call
  • Redirecting it
  • Sending it to a voice CAPTCHA
  • Tagging it for the agent
  • Routing it to a trained team
  • Alerting security operations
  • Triggering additional verification
  • Preventing an associated outbound call

5. Learn

This may be the most important capability.

Threat intelligence cannot remain static.

Attackers change telephone numbers.

They change calling behavior.

They change scripts.

They change targets.

They change techniques.

A modern voice security system therefore needs a feedback loop:

Detect → Analyze → Respond → Confirm → Learn → Detect again

That is how voice security moves from static blocking to adaptive defense.


Why Static Blocklists Are Failing

A traditional approach to unwanted calls is simple:

Find bad numbers. Block bad numbers.

The problem is that attackers can change numbers.

A static blocklist is therefore useful, but insufficient.

Consider an IVR callback scam.

An attacker can use different numbers, change destinations or modify the behavior of the call.

Assertion’s research on IVR callback scams highlights why static blocklisting and manual intervention struggle against dynamic attack patterns, and why automated screening combined with continuously updated threat intelligence is more effective.

The same principle applies to broader voice threats.

The threat intelligence needs to move as quickly as the threat.


Contact Centers Are Becoming a High-Value Voice Target

Contact centers are uniquely exposed.

They combine:

High call volumes + human agents + customer information + business authority.

An attacker does not necessarily need to compromise the contact-center platform.

They may only need to convince an agent.

That makes contact-center security fundamentally different from simply securing the underlying telecom infrastructure.

A contact center needs to know:

  • Who is calling?
  • Has this number appeared before?
  • What has this caller done previously?
  • Is the calling behavior abnormal?
  • Is this caller associated with known threats?
  • Is this interaction consistent with normal customer behavior?
  • Should the call reach a general agent?
  • Should it go to a specialist?
  • Should the caller be challenged?
  • Should the call be blocked?

This is voice threat detection at the point of interaction.


TDoS: When the Attack Is Not About Fraud

Not every voice attack is designed to steal information.

Sometimes the objective is simply to make the telephone unavailable.

This is Telephony Denial of Service (TDoS).

An attacker can generate large volumes of calls or signaling activity to overwhelm a telephone number, IVR, contact center or emergency service.

The result can be operational disruption.

Legitimate callers cannot get through.

Agents become overloaded.

Critical calls can be delayed.

A recent Assertion analysis describes how TDoS can overwhelm healthcare lines and how behavioral filtering and secondary screening can help distinguish legitimate callers from attack traffic.

This is why voice availability is also a security issue.

For a contact center, hospital or financial institution:

An unavailable telephone system can become a business continuity incident.


Toll Fraud Is a Voice Security Problem Too

Another category is toll fraud.

In a toll-fraud attack, an attacker abuses the organization’s voice infrastructure to generate unauthorized or expensive calls.

The organization may discover the problem only after the telecom bill arrives.

That is too late.

The better approach is to identify suspicious outbound behavior in real time.

This can include:

  • Unusual international destinations
  • Premium-rate destinations
  • Sudden changes in calling volume
  • Abnormal call duration
  • Unauthorized destinations
  • Calls outside established policies

Voice security therefore has to cover both sides of the conversation:

Inbound security.

and

Outbound security.

Assertion’s Secure Voice platform is designed around this broader model, screening incoming and outgoing calls and applying configurable security policies.


IVR Systems Are Also an Attack Surface

IVRs are designed to make customer service more efficient.

But they can also become a target.

Attackers can repeatedly probe IVR menus.

They can exploit callback functionality.

They can test authentication flows.

They can generate abnormal call patterns.

They can attempt to discover valid extensions.

They can use repeated interactions to learn how an organization behaves.

Assertion’s research on IVR account-access scams describes how behavioral signals such as repeated disconnects, short-duration calls and unusual interaction patterns can help identify suspicious activity before the attacker progresses further.

This leads to an important principle:

The IVR should not be treated as a neutral front door. It is part of the security perimeter.


Healthcare Has a Different Voice Security Problem

Healthcare organizations have an unusually high dependence on voice.

Patients call hospitals.

Families call patient rooms.

Clinicians communicate through voice systems.

Insurance members call contact centers.

Healthcare organizations also operate large, distributed telephony environments.

That creates a broad attack surface.

Assertion’s healthcare-specific Secure Voice offering identifies several risks, including patient-room scam calls, robocalls, spoofing and the challenge of identifying CMS-related auditor calls within ordinary inbound traffic.

The operational impact is significant.

A nuisance call in a normal enterprise is frustrating.

A nuisance call into a patient’s room is different.

It can interrupt care.

It can create confusion.

It can expose vulnerable patients to scams.

And it can consume staff time.

Healthcare therefore needs to think about voice security not just as fraud prevention, but also as patient safety and operational protection.


The CMS Auditor Problem: When the Important Call Looks Like Every Other Call

Healthcare insurers and organizations can also face another voice challenge:

How do you identify a high-value inbound call when the caller is intentionally trying not to be identified?

Assertion’s healthcare research focuses on CMS auditor call detection as one such use case.

The core challenge is that auditor calls can resemble ordinary customer-service traffic, while the organizations receiving them need to handle them appropriately.

The underlying model is particularly interesting from a voice intelligence perspective.

A static list is not enough.

Numbers can change.

Therefore, the intelligence needs to evolve.

Assertion’s approach described in the Voice Threat Analysis source uses a shared, continuously updated database in which reported intelligence is aggregated, reconciled and redistributed to participating customer environments. The system also incorporates feedback so that confirmed false positives can reduce confidence in future classifications.

This illustrates a broader principle:

Voice intelligence gets stronger when every interaction can contribute to the next decision.


The Architecture of Modern Voice Security

The architecture of enterprise voice security is evolving.

A simplified model looks like this:

Caller

Carrier / Network

SBC

Voice Security Layer

IVR / Contact Center / UC

Agent

The key difference is what happens between the SBC and the human.

Instead of allowing every call to proceed based primarily on network-level controls, the voice security layer can evaluate the interaction before it reaches the agent.

This creates an opportunity to stop threats earlier.

And earlier is better.

Because once a malicious caller is speaking to an employee, the security system is no longer protecting just infrastructure.

It is trying to protect a human being from manipulation.


What Should a Modern Voice Security System Detect?

A mature voice security platform should look beyond caller ID.

It should consider signals such as:

Caller intelligence

Has the number been associated with suspicious behavior?

Calling behavior

Is the pattern consistent with normal customer behavior?

Frequency

Is the caller generating unusual call volumes?

Duration

Are there repeated extremely short or unusually patterned calls?

Signaling

Are there network or SIP indicators associated with abnormal behavior?

Destination

Is the caller targeting unusual numbers, departments or extensions?

History

Has this caller interacted with the enterprise before?

Threat intelligence

Is the number or behavior associated with known attacks?

Context

Does the interaction make sense for the organization, geography, time and service being accessed?

The result should not be a simplistic yes/no answer.

It should be a risk decision.


Voice Security Should Be Risk-Based, Not Rule-Based

Rules still matter.

Geo-fencing matters.

Blocklists matter.

OFAC policies matter.

Time-of-day policies matter.

Calling thresholds matter.

But rules alone cannot keep up with a dynamic threat landscape.

The better model is:

Rules + intelligence + behavior + AI + human feedback

This enables the system to distinguish between:

Known threat

Suspicious behavior

Unknown caller

Legitimate caller

That distinction reduces unnecessary blocking while increasing protection against emerging threats.


What Happens When a Suspicious Call Is Detected?

Detection is only useful if the organization can act.

A modern system should give the enterprise multiple response options.

Block

For calls with sufficiently high confidence of malicious intent.

Challenge

For suspicious or first-seen callers where more evidence is required.

Tag

Warn the agent before or as the call reaches them.

Route

Send suspicious interactions to specially trained agents.

Alert

Notify supervisors or security teams.

Monitor

Allow the interaction while increasing visibility.

Learn

Feed the result back into the intelligence layer.

This is particularly important in contact centers.

A suspicious call does not always need to be blocked.

Sometimes the better security decision is:

“Let a trained person handle this.”


The Goal Is Not to Block More Calls

This is where voice security differs from traditional spam blocking.

The objective is not:

Block as many calls as possible.

The objective is:

Block the calls that should not be there while protecting the calls that matter.

That means voice security has to optimize for both:

Security

and

Availability.

A healthcare organization cannot simply block unknown callers.

A bank cannot block every first-time customer.

A contact center cannot reject every unusual interaction.

The system needs to understand context.

That is why behavioral intelligence is so important.


From Voice Monitoring to Voice Security

There is a fundamental difference between these two approaches.

Voice monitoring

“What happened?”

Voice security

“What is happening, how risky is it, and what should we do?”

Monitoring produces logs.

Security produces decisions.

Monitoring tells you a call was made.

Security can determine whether that call should be allowed, challenged, blocked, routed or escalated.

This is the shift enterprises need to make.


Why Voice Security Belongs in the CISO Conversation

Historically, voice belonged to telecom.

Security belonged to IT.

That division is becoming increasingly artificial.

The voice system can now be:

  • An entry point for attackers
  • A fraud channel
  • A social-engineering channel
  • A business-continuity dependency
  • A source of sensitive information
  • An attack target
  • A mechanism for financial loss

That makes voice security relevant to:

CISOs

Chief Information Officers

Voice Operations leaders

Contact Center leaders

Fraud teams

Risk teams

Compliance teams

Healthcare security teams

The ownership may still sit with different departments.

The risk does not.


What Should CISOs Ask About Voice Security?

A simple security review can begin with ten questions:

  1. Can we detect suspicious calls before they reach our agents?
  2. Can we identify voice threats in real time?
  3. Can we detect abnormal calling behavior?
  4. Can we challenge suspicious callers?
  5. Can we route high-risk calls to trained agents?
  6. Can we protect our IVRs from abuse?
  7. Can we detect TDoS activity before it overwhelms operations?
  8. Can we detect toll-fraud activity on outbound calls?
  9. Does our voice threat intelligence continuously learn from new attacks?
  10. Can our voice security layer integrate with our existing SBC, contact center and UC infrastructure?

If the answer to several of these is no, there is probably a voice security gap worth investigating.


How to Build an Enterprise Voice Security Strategy

A practical strategy can be built in six layers.

Layer 1: Establish Visibility

Know what enters and leaves the voice environment.

You cannot secure what you cannot see.

Layer 2: Establish Intelligence

Build access to threat intelligence and behavioral data.

Layer 3: Establish Risk

Score calls based on multiple signals instead of relying on a single rule.

Layer 4: Establish Controls

Define what happens at each risk level.

Layer 5: Establish Human Escalation

Give supervisors and trained agents the context required to handle suspicious interactions.

Layer 6: Establish Continuous Learning

Use confirmed threats and false positives to improve future decisions.

This creates a voice-security operating model rather than another point security product.


Where Assertion Secure Voice Fits

This is the problem Assertion Secure Voice is built to address.

Secure Voice sits alongside the existing voice infrastructure rather than requiring an organization to replace its carrier, SBC or contact-center platform. Assertion describes the solution as an AI-enabled layer that screens incoming and outgoing calls, detects anomalies and applies controls in real time.

Its capabilities include:

  • AI-based call screening
  • Incoming and outgoing call analysis
  • Scam and robocall detection
  • Toll-fraud protection
  • TDoS protection
  • Voice CAPTCHA
  • Suspicious-call tagging and routing
  • Remote-worker SBC protection
  • Threat intelligence
  • Custom security policies
  • Healthcare-specific voice security use cases

The architectural principle is simple:

Do not replace the voice infrastructure. Add intelligence and security around it.


Why the Future of Voice Security Is Adaptive

Attackers evolve.

Therefore, voice security must evolve.

A security platform that depends entirely on yesterday’s blocklist will struggle with tomorrow’s attack.

A platform that continuously combines:

Threat intelligence

Behavioral analysis

AI

Network signals

Human feedback

can become progressively better at identifying suspicious interactions.

That is the foundation of adaptive voice security.


The Phone Did Not Suddenly Become Dangerous

This is perhaps the most important lesson.

The telephone was not magically transformed into an attack surface.

It always was one.

We simply treated it differently.

Email became a security problem.

We built email security.

Endpoints became a security problem.

We built endpoint security.

Identity became a security problem.

We built identity security.

Cloud became a security problem.

We built cloud security.

Now voice is becoming a security problem.

The security architecture needs to catch up.


The Next Security Perimeter Is Already Ringing

The next major voice-security incident may not begin with a sophisticated exploit.

It may begin with a phone call.

A caller who sounds legitimate.

A patient who should never have been contacted.

An agent who receives a convincing request.

An IVR that is quietly being probed.

A contact center being flooded with malicious traffic.

An employee being targeted by an AI-generated voice.

A number that looks familiar.

A conversation that sounds normal.

Until it isn’t.

That is why enterprise voice security can no longer be treated as simply a telecom problem.

Voice is now a security surface.

And every enterprise that depends on the telephone should know what is happening on it.


The Voice Security Checklist

Before you consider your enterprise voice environment protected, ask:

  • Do we screen inbound calls in real time?
  • Do we screen outbound calls for fraud and policy violations?
  • Can we identify suspicious callers before they reach agents?
  • Do we use behavioral analysis rather than static blocklists alone?
  • Can we challenge first-seen or suspicious callers?
  • Can we detect TDoS patterns?
  • Can we detect toll-fraud activity?
  • Can we protect IVRs from abuse?
  • Can we protect remote-worker SBCs?
  • Can we route suspicious calls to trained agents?
  • Can supervisors receive real-time threat alerts?
  • Can confirmed threats update our intelligence?
  • Can false positives improve future decisions?
  • Does voice security integrate with our existing voice infrastructure?
  • Does our security team have visibility into voice threats?

If you cannot confidently answer yes to most of these, your voice infrastructure deserves a security assessment.


Frequently Asked Questions About Voice Security

What is enterprise voice security?

Enterprise voice security protects corporate phone systems, contact centers, SIP infrastructure and voice applications against threats including vishing, robocalls, scam calls, toll fraud, TDoS, impersonation and other malicious voice activity. Unlike basic call blocking, enterprise voice security combines threat intelligence, behavioral analysis, AI and real-time controls.

Why is voice security important?

Voice remains a critical channel for customer service, healthcare, banking, insurance and business operations. Attackers can exploit that trust through social engineering, fraud, impersonation, denial-of-service attacks and increasingly AI-generated voice interactions.

What is voice threat detection?

Voice threat detection is the process of identifying suspicious or malicious calls using signals such as caller intelligence, network information, calling behavior, historical activity, threat intelligence and contextual risk.

What is vishing?

Vishing, or voice phishing, is a social-engineering technique in which attackers use telephone calls to deceive victims into revealing information, transferring money, approving access or performing another action. CISA categorizes spearphishing via voice as an Initial Access technique.

Can STIR/SHAKEN stop vishing?

No. STIR/SHAKEN helps authenticate caller identity information, but authentication does not establish that the caller’s intent is legitimate. Call analytics and additional security controls are still required to assess risk.

What is TDoS?

TDoS stands for Telephony Denial of Service. It involves overwhelming a telephone service, number, IVR or contact center with malicious traffic so legitimate callers have difficulty getting through.

What is toll fraud?

Toll fraud is the unauthorized use of an organization’s voice infrastructure to make calls that can generate significant telecom charges. Real-time monitoring and outbound calling policies can help detect and prevent suspicious activity.

Why aren’t SBCs enough for voice security?

SBCs provide important network-edge controls, but modern voice threats also require behavioral analysis, threat intelligence and real-time risk decisions. Voice security adds an intelligence layer focused specifically on the behavior and risk of calls.

How does AI improve voice security?

AI can analyze large volumes of calling behavior and identify patterns that are difficult to detect using static rules alone. It can help identify anomalies, score suspicious interactions and continuously improve detection using feedback.

What is Zero Trust Voice Security?

Zero Trust Voice Security applies the principle of continuous verification to telephone interactions. Instead of assuming a caller is trustworthy because the number appears legitimate, the system evaluates multiple signals before deciding how the interaction should be handled.

Can voice security protect contact centers?

Yes. Voice security can screen calls before they reach agents, identify suspicious behavior, block malicious calls, challenge uncertain callers, route high-risk calls to trained agents and provide security teams with real-time visibility.

Can voice security protect healthcare organizations?

Yes. Healthcare organizations can use voice security to reduce patient-room spam and scam calls, protect contact centers, detect suspicious calling patterns and support specialized workflows such as CMS-related auditor call detection. Assertion offers a healthcare-specific Secure Voice solution.

What should a CISO look for in a voice security platform?

A CISO should evaluate real-time call screening, behavioral analytics, threat intelligence, AI-based detection, TDoS protection, toll-fraud controls, IVR security, caller challenges, policy enforcement, alerting, learning capabilities and integration with existing SBC and contact-center infrastructure.


The Definitive Resource for Enterprise Voice Security

Voice is no longer just infrastructure.

It is a business channel.

It is a customer channel.

It is a revenue channel.

And increasingly, it is an attack channel.

The organizations that treat it as a security perimeter will be better positioned to detect threats before they become incidents.

Because the next attack may not arrive in your inbox.

It may not exploit your endpoint.

It may not trigger an identity alert.

It may simply ring.

Download the Voice Threat Analysis 2026

Read Assertion’s research into the changing voice threat landscape, the rise of vishing and contact-center fraud, and the security implications for organizations in healthcare, banking, contact centers and government.

[Download Voice Threat Analysis 2026]

See Enterprise Voice Security in Action

Explore how Assertion Secure Voice can screen calls, detect voice threats, apply real-time controls and protect enterprise voice infrastructure without replacing your existing carrier, SBC or contact-center platform.

[Explore Assertion Secure Voice]


Continue Reading: The Voice Security Knowledge Hub

Build your organization’s voice-security strategy with these related resources: