
Closed
Posted
Paid on delivery
Project Description: Title: Urgent Support Required: Skype for Business 2019 Enterprise Pool - Windows Fabric Quorum Loss Recovery Hello Everyone, We are facing a critical production issue with our Skype for Business Server 2019 Enterprise Pool environment consisting of 3 Front End Nodes, and we need an expert to help us bring it back online safely. Current Situation & Backend Infrastructure Status: SQL Backend: The backend is running on SQL Server AlwaysOn. Following a major database outage, the original Availability Group Listener configuration was corrupted/deleted. Listener Restored: We have already fixed the database layer. We successfully re-created the AG Listener via T-SQL under the exact name expected by the topology (sfb-sql-listene) on a new IP ([login to view URL]). The listener is now completely Online inside the Windows Failover Cluster Manager, and the corresponding A-record has been fully updated in our Infoblox DNS. The Core Issue: The Skype for Business Front End services (RTCSRV and RTCCPS) fail to start. Because the entire 3-node pool dropped simultaneously during the initial SQL crash, we are stuck in a hard Windows Fabric Quorum Loss state. The local services clear their cache but drop back to a "Stopped" state upon manual service startup attempts due to the lack of a cluster majority. What we need from you: We need an expert to safely guide our engineering team through recovering the Windows Fabric cluster and forcing/syncing the core Front End services back to a Running state. We prefer to execute one of the following approaches based on your assessment: Properly force-start the pool on the primary node by bypassing the standard quorum constraint via management shell (Start-CsPool -IsPrimaryComputerFailureRecovery). Or coordinate and sync the clean initial startup sequence between the multiple Front End nodes (FE-1 and FE-2) to satisfy the native Windows Fabric quorum safely without causing a split-brain scenario. Requirements: Proven hands-on experience with Skype for Business Server 2019 Enterprise Pools. Deep expertise in Windows Fabric Clustering troubleshooting and lifecycle management. Strong knowledge of SQL AlwaysOn integration with SfB Topologies. Please apply only if you have resolved exact Windows Fabric quorum loss scenarios in enterprise environments before. Ready to start immediately.
Project ID: 40531565
17 proposals
Remote project
Active 6 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
17 freelancers are bidding on average $455 USD for this job

Hello, I've worked on Microsoft unified communications environments involving Skype for Business Enterprise Pools, SQL AlwaysOn backends, Windows Failover Clustering, and service recovery after infrastructure failures. The challenge is recovering Windows Fabric and re-establishing Front End quorum so RTCSRV and RTCCPS can successfully join the cluster again. My recovery process would focus on: • Verifying FabricHost, Fabric membership, and replica health on all FE nodes • Confirming connectivity from each Front End to the restored AG Listener • Reviewing Fabric and Skype logs for quorum-related failures • Identifying whether the pool requires primary failure recovery or coordinated node synchronization • Safely bringing services back online while preventing split-brain conditions • Validating replication, CMS access, and Front End service health after recovery I've handled situations where an outage left all Front End servers offline simultaneously, causing Fabric to lose cluster majority and preventing normal service startup even after SQL was repaired. In those cases, the key is restoring Fabric state correctly rather than repeatedly attempting to start services. Once recovered, I'll help verify that the pool is stable, services remain running, and user-facing functionality is restored before considering the incident closed. I can start immediately and work alongside your engineers during the recovery window. Regards, Chaz C.
$750 USD in 2 days
2.4
2.4

I can quickly assist you in recovering your Skype for Business Server 2019 setup. With my solid background in SQL and troubleshooting, I can effectively guide your team through the recovery process of the Windows Fabric cluster. My hands-on experience with SQL Server AlwaysOn and deep knowledge of the Skype for Business architecture will ensure a smooth recovery. I understand the urgency, and I’m ready to jump in immediately to help navigate through the quorum loss state. What specific recovery actions have been attempted so far? Would you prefer to force-start the pool on the primary node, or sync the startup sequence between the nodes? Have you attempted any specific recovery steps before this? What exact issues have you faced while trying to restart the RTCSRV and RTCCPS services?
$250 USD in 11 days
1.5
1.5

Hello, I can help recover your Skype for Business 2019 Enterprise Pool and troubleshoot the Windows Fabric quorum loss safely. Based on your details, SQL AlwaysOn and the listener look restored, so the key risk is bringing RTCSRV/RTCCPS back without split-brain or stale Fabric state. I’ll first validate topology, SQL listener connectivity, Fabric health, event logs, FE node status, DNS resolution, and service dependencies. I’ll then guide your team through the safest recovery path, either controlled Start-CsPool -IsPrimaryComputerFailureRecovery on the right primary node or a coordinated multi-node startup sequence if quorum can be restored cleanly. I have strong experience with SQL Server, clustering, production outage recovery, and enterprise infrastructure troubleshooting, and I’m comfortable working live with engineering teams during critical incidents. I am ready to start immediately. Best regards, Smit
$250 USD in 3 days
1.6
1.6

I looked at your Skype for Business 2019 Enterprise Pool situation and see you have a 3-node Windows Fabric cluster stuck in quorum loss after a SQL AlwaysOn listener outage, with RTCSRV and RTCCPS services failing to start despite the backend being restored. The key to resolving this is using the `Start-CsPool -PoolFqdn <FQDN> -QuorumLossRecovery` cmdlet, which forces the pool to start by reloading user data from the backup store for routing groups in quorum loss . For a 3-node pool, the first startup requires all three servers running, but subsequent starts can recover with fewer if you use the `Reset-CsPoolRegistrarState -ResetType QuorumLossRecovery` cmdlet before bringing the pool back up . I have resolved exact Windows Fabric quorum loss scenarios in enterprise environments and can guide your team through the recovery sequence to avoid split-brain. The project will be completed in 1 day at a total cost of 300 USD. Do you have the output of `Get-CsPoolFabricState -PoolFqdn <FQDN> -Type Routing` available to identify which routing groups are affected? Should the recovery use `-QuorumLossRecovery` with the main pool or consider skipping specific routing groups with `-SkipRoutingGroup` if some are unrecoverable? Let me know your answers. I can start right away.
$300 USD in 1 day
0.0
0.0

Hi, I have extensive experience with Skype for Business 2019 Enterprise Pools and Windows Fabric clustering. I’ve successfully recovered pools from hard quorum loss scenarios and coordinated SQL AlwaysOn integration with SfB topologies in production environments. My approach would be: ✓ Assess the current Fabric quorum and Front End service state ✓ Safely force-start the primary node using Start-CsPool or coordinate multi-node startup ✓ Validate RTCSRV and RTCCPS services are fully running without risk of split-brain ✓ Ensure the listener and DNS configuration remain consistent One quick question: Are you looking for a full hand-holding live session with your team, or a step-by-step remote guide for execution? I can start immediately and work carefully to restore your pool safely. Regards, Sergio
$250 USD in 1 day
0.0
0.0

Hello, Based on your description, Front End pool is still trapped in a Windows Fabric quorum-loss condition after all three FE nodes dropped simultaneously during the database outage. At that point, restoring the AG Listener alone is not enough because RTCSRV and RTCCPS will continue failing until Fabric membership and quorum are re-established correctly. My approach would be to first validate the current Fabric state across all FE servers, confirm cluster membership, replica status, FabricHost health, and verify that the restored AG Listener is fully reachable from each Front End node using the topology's expected configuration. From there, I would determine whether the safest recovery path is: • Controlled pool recovery using Start-CsPool -IsPrimaryComputerFailureRecovery • Coordinated Front End startup sequence to re-establish Fabric quorum • Additional Fabric cleanup/synchronization if stale cluster metadata remains after the outage I have experience troubleshooting Skype for Business Enterprise environments, SQL AlwaysOn integrations, Windows Failover Clustering, service startup failures, and post-disaster recovery scenarios where Fabric quorum loss prevents RTCSRV/RTCCPS from coming online. Before making changes, I would verify the topology state and Fabric health to avoid creating a split-brain condition or introducing further cluster inconsistency. I am available to start immediately and work directly with your engineering team during recovery. Regards,Remon
$750 USD in 3 days
0.0
0.0

Hello, I can help with "Skype for Business Server Recovery Expert" and keep the work clean and practical. My focus would be making the backend and connected services work smoothly together. The listed skills point to SQL Troubleshooting Technical Support Database Administration Network Administration Microsoft SQL Server System Administration, so I would keep the implementation aligned with that. I would first check the current setup, then complete the work in a way that is easy to review. Quick infrastructure questions: 1. Which user roles, database records, and admin actions need to be supported? 2. Should the solution be optimized for future scaling, easier maintenance, or a simple handover? 3. Which external services, APIs, or payment gateways need to be connected, and will test credentials be available? Regards, Houssame
$500 USD in 7 days
0.0
0.0

Alhasa, Saudi Arabia
Payment method verified
Member since Feb 16, 2020
$30-250 USD
$30-250 USD
$30-250 USD
$30-250 USD
₹12500-37500 INR
$3-10 NZD / hour
$25-50 USD / hour
$10-30 USD
$25-50 USD / hour
$250-750 NZD
$10-30 USD
$100-300 USD
$750-1500 USD
$30-250 USD
$1500-3000 USD
₹12500-37500 INR
$100-225 USD / hour
₹400-750 INR / hour
$250-750 AUD
$250-750 USD
$250-750 USD
$8-15 CAD / hour
£750-1500 GBP
₹37500-75000 INR