AZ-500 - Securing an Azure Hub-Spoke Network
📌 Overview
This project applies the Secure Networking domain of the AZ-500 (Azure Security Engineer Associate) exam — at 20–25%, the largest single domain on the exam — to a real environment rather than a standalone lab. The hub-spoke topology itself isn't new: it's the same VNets, VPN Gateway, and hybrid connectivity built for the AZ-700 project, renamed and reused. What's new is everything layered on top of it: NSGs and ASGs, a centralized Azure Firewall, user-defined routes, and a Network Watcher validation pass that ended up being the most instructive part of the whole build.
Environment:
- Hub VNet (JPVNetHub) with a Virtual Network Gateway and a centralized Azure Firewall
- Two spoke VNets (JPVNetSpoke1, JPVNetSpoke2) peered to the hub, each running one workload VM
- Site-to-Site VPN to an on-prem FortiGate, Point-to-Site VPN authenticated through Microsoft Entra ID
- Three VMs (JPAZVM11 hub, JPAZVM12 Spoke1, JPAZVM13 Spoke2), each behind its own NSG and Application Security Group
🔧 Objectives
- Apply defense-in-depth network controls — NSGs/ASGs, centralized firewall, UDRs — to an existing topology instead of a disposable lab
- Get hands-on with the specific AZ-500 Secure Networking skills: segmentation, egress control, VPN connectivity, and Network Watcher diagnostics
- Validate every control live rather than assuming configuration equals enforcement
- Produce a portfolio-ready case study, including the troubleshooting that didn't go as expected
🔒 Segmentation: NSGs & ASGs
The first surprise of this project showed up before any new resources were even deployed: peering two VNets together does not segment them. Azure's platform-default AllowVnetInBound rule (priority 65000) quietly allows everything inside the peered address space, and it can't be deleted — only out-prioritized. Phase 1 added an explicit Deny-VirtualNetwork-Inbound rule (priority 4000) and Deny-Internet-Inbound (4010) to all three NSGs, then layered narrow allows above them for exactly what's needed: RDP and ICMP from on-prem, plus RDP from the hub as a jump-host path.
Application Security Groups (asg-hub-infra, asg-spoke1-workload, asg-spoke2-workload) sit between the NSG rules and the VMs so rules reference a role instead of a hardcoded IP. With one VM per role today that's not saving much, but it means the rules don't need to change if a role ever grows past one VM.
🔥 Centralized Firewall & Egress Control
Segmentation handles north-south trust at the NSG layer, but spoke-to-spoke traffic over VNet peering has no inspection point in the path at all by default — peering just connects, it doesn't route through anything. Phase 2 deployed an Azure Firewall (Standard SKU) into the hub and built a policy with a network rule permitting HTTPS egress from both spokes and an application rule allowing Windows Update by FQDN tag, with Threat Intelligence in Alert mode to start.
⚠️ The UDR Precedence Gotcha
This phase had its own gotcha, and it's the one worth remembering: Azure resolves routing by longest-prefix match, and a user-defined route only beats a system route at equal or greater specificity. The automatic peering route between the two spokes and the explicit UDR pointing spoke-to-spoke traffic at the firewall are both /16 prefixes — so adding only a 0.0.0.0/0 → Firewall default route wasn't enough. Without an explicit 10.2.0.0/16 → Firewall route on Spoke1's route table (and the mirror on Spoke2's), the automatic peering route kept winning and spoke-to-spoke traffic silently bypassed the firewall entirely. Nothing errored. Nothing warned about it. The only way to catch it was to check Next Hop and see VirtualNetwork instead of VirtualAppliance — which is exactly what Phase 4 is for.
🔐 Secure VPN Connectivity
No new build here — the Site-to-Site (custom IPsec/IKE policy, AES256/SHA256/DH Group 14/PFS2048) and Point-to-Site (Entra ID authentication, OpenVPN) connections already existed from the AZ-700 project. This phase was about reviewing what was already there and mapping it explicitly to this domain's requirements: a custom cipher suite instead of accepting Azure's default proposal, and identity-based P2S auth instead of a shared certificate.
🔍 Visibility & Validation
This phase deploys nothing — it's a validation pass against the earlier phases, run through five Network Watcher tools: Effective Security Rules, Next Hop, IP Flow Verify, NSG Diagnostics, and Connection Troubleshoot.
Flow logs were evaluated for this phase first, and dropped. Both NSG Flow Logs and Virtual Network Flow Logs produce deeply nested JSON with no simple allow/deny field — flow state gets folded into a single value instead — and critically, they never capture the IP Flow Verify or NSG Diagnostics tests, since neither of those tools sends a real packet; they're rule-evaluation simulations. Network Watcher's other diagnostics validate the same enforcement more directly and read far better live, so flow logging was left out of the build rather than bolted on as a sixth method.
🚧 The Three-Layer RDP Mystery
The last test in the validation plan was Connection Troubleshoot — real traffic, not a simulation — from JPAZVM12 (Spoke1) to JPAZVM13 (Spoke2) on port 3389. By design this was expected to come back Unreachable, since Spoke2's NSG only allows RDP from the hub subnet and on-prem, not from Spoke1 directly. It did:
Connection Troubleshoot: Unreachable, 316/316 probes failed — the expected result. (Click to enlarge.)
That was the expected result — right up until I modified Allow-RDP-From-Hub to add JPAZVM12's IP as an explicit source, intending to test opening that one host through. The NSG rule now clearly allowed it. Connection Troubleshoot still said Unreachable, 316/316 failed.
Troubleshooting approach: work outward from the NSG, since that's the layer I'd just changed. First, confirm outbound wasn't the problem — the Spoke1 NSG's outbound rules were untouched, just the platform defaults:
Spoke1's outbound rules, unmodified — the default AllowVnetOutBound already covers this. (Click to enlarge.)
The VirtualNetwork service tag includes peered VNets, so outbound from Spoke1 to Spoke2 was never the issue. With NSG outbound on Spoke1 allowed, NSG inbound on Spoke2 allowed (after my edit), and still 100% probe failure, there was exactly one checkpoint left on the path: Azure Firewall. UDRs route all Spoke1↔Spoke2 traffic through the firewall, and its policy only had a network rule for TCP 443 — nothing for 3389. Azure Firewall default-denies anything with no matching rule, so it was silently dropping every SYN regardless of what either NSG said.
Adding a network rule fixed it — permitting TCP 3389 from Spoke1 to Spoke2:
The missing firewall rule, added — TCP 3389, Spoke1 to Spoke2, Allow. (Click to enlarge.)
Re-running Connection Troubleshoot: Reachable, 316/316 probes passed.
Reachable — once both the NSG and the firewall agreed. (Click to enlarge.)
Except the full diagnostic (Connectivity + NSG diagnostic + Next Hop + Port Scanner, all in one pass) told a more complete story — three checks passed clean, and a fourth came back stuck:
Everything upstream passes — but "Destination port accessible" times out. (Click to enlarge.)
"Destination port accessible" is a different question than "can a packet get there" — it's asking whether something is actually listening on 3389 from the destination VM's own perspective. A timeout there, with everything upstream green, points past Azure entirely and into the guest OS. Sure enough: these VMs had never been RDP'd into, so Windows Defender Firewall on JPAZVM13 was still blocking RDP at its default, out-of-the-box state — a layer Network Watcher can diagnose the symptom of but can't reach to fix.
That's three independent checkpoints in one troubleshooting pass — NSG, Azure Firewall, and guest OS firewall — each one caught something real, each one needed a different tool and a different fix. I scripted the last one using Invoke-AzVMRunCommand so it can flip fDenyTSConnections and enable the Windows Firewall's Remote Desktop rule group over the control plane, without needing RDP to already work to get in and fix it:
Set-ItemProperty -Path "HKLM:\System\CurrentControlSet\Control\Terminal Server" `
-Name "fDenyTSConnections" -Value 0
Enable-NetFirewallRule -DisplayGroup "Remote Desktop"
📈 Results
- Earlier phases confirmed enforcing via live diagnostics, not just reviewed configuration — Effective Security Rules, Next Hop, IP Flow Verify, and NSG Diagnostics all returned expected results
- Traced a real "it should work but doesn't" failure through three independent enforcement layers (NSG → Azure Firewall → guest OS) to its actual root cause at each step, rather than guessing
- Flow logs evaluated and deliberately left out of the final build, in favor of Network Watcher's diagnostic tools, which cover the same ground and read far better live
- VPN connectivity from the AZ-700 build reviewed and explicitly mapped to AZ-500's Secure Networking requirements, with no changes needed
- A reusable Run Command script added to the project for enabling RDP on fresh VMs ahead of a live demo, without needing the network path open first
📝 Notes / Lessons Learned
- Passing one layer doesn't mean the request passed every layer — an NSG allow, a UDR pointing at the firewall, and the firewall's own rule set are each an independent decision point. "I allowed it" only answers one of those three questions at a time
- Azure Firewall fails closed and silently — with no matching rule, it just drops the packet, no NSG-style "Deny" result to point at. The only visible symptom from the client side is a timeout indistinguishable from a dozen other causes, which is why a tool that sends real traffic is worth running even after the simulated tools say everything's fine
- The UDR precedence issue is the kind of bug that doesn't announce itself — two routes at the same prefix length, one silently winning over the other, with no error on either side. Worth validating routing explicitly instead of assuming a UDR took effect just because it exists
- Raw flow log JSON looks like ground truth and isn't, for this use case — it's a legitimate record of some traffic, but the specific tools used to validate enforcement here never touch the wire, so the tests this project cared about never would have shown up in it regardless of how well-parsed the output was