Tesla Referral P40 • 48GB Installed GPUMemory • 22TB+ Storage • 25GbE + 40Gb Networking
• A nonprofit funding / grants / communications environment
• Both Tesla Referral P40 GPUs actively processing AI workloads
• Real AI interaction with operational data
• Funding-opportunity and requirement analysis
• AI-generated Word / Excel / PDF documents
• Multiple documents assembled into a single PDF package
• Incoming reply synchronization and AI interpretation
• Human-reviewed outbound response preparation
• Multi-user generated workspace output
• Local ~32B quantized language model for reasoning
• Separate 7B vision model for image understanding
• Custom local API / bridge
• Agent / tool execution loop
• 65 locally exposed tools
• Open WebUI as an alternate interface
• GPU-backed image generation
• OCR and document-processing capabilities
↓
AI securely queries operational data using read-only access
↓
Analyzes the returned information
↓
Summarizes the findings
↓
Creates a Word or PDF report if requested
↓
Returns the finished result to the user
↓
Vision + OCR
↓
AI Interpretation
↓
Structured Data
↓
Excel / Word / PDF Output
↓
Opportunity and funder organized in the portal
↓
Eligibility rules, restrictions, geographic requirements, deadlines, andsupporting information analyzed
↓
Requirements compared against the organization's own knowledge and records
↓
AI evaluates whether the opportunity is worth pursuing and explains why
↓
Correspondence drafted
↓
Supporting Word / Excel / PDF documents generated
↓
Multiple documents assembled into one PDF package when required
↓
Email body and attachments prepared
↓
↓
Incoming reply synchronized back into the workflow
↓
AI reads the reply in context with the prior correspondence and funder'srequirements
↓
Next response prepared for human approval
• File creation, reading, updating, and organization
• Microsoft Word generation
• Excel generation
• PowerPoint generation
• PDF generation
• PDF merging and document-package assembly
• OCR
• Image understanding
• GPU-backed image generation
• Media processing
• Controlled Windows / PowerShell automation
• Research and external data ingestion
• Automatic startup
• Boot-delay handling
• Watchdog / automatic service restart
• Persistent AI memory
• Multi-user workspaces
• Per-user folders
• Real generated output
• Runtime logging
• Restricted filesystem paths
• LAN-only service restrictions
• Dedicated least-privilege database account
• Read-only database access
• Explicit restriction from unrelated databases
• No database administrator rights
• Bare enterprise servers that still need GPUs, drivers, software, andconfiguration
• High-density GPU systems requiring much greater power, cooling, andsupporting infrastructure
288GB ECC system memory
2× NVIDIA Tesla Referral P40 GPUs with 24GB each
22TB+ installed storage
4TB raw NVMe capacity
25GbE + 40Gb high-speed networking
iDRAC9 Enterprise
Dual redundant 1100W power supplies
• 24 cores / 48 threads each
• 48 total cores / 96 total threads
• 10 DIMMs installed
Tesla Referral P40
• 24GB VRAM per GPU
• 48GB total installed GPU memory across two discrete GPUs
• Headless enterprise compute GPUs with no display outputs
• Approximately 17.46TB RAID data volume
• Approximately 1.09TB RAID volume
• 4× 1TB NVMe SSDs
• Approximately 4TB raw NVMe capacity
• 2× Mellanox ConnectX-3 40G supporting IB/IPoIB
• Dell Y3H8J wide-range AC units
• Remote power control
• Hardware monitoring
• Temperature monitoring
• Firmware management
• System-health monitoring
• Out-of-band administration
Tesla Referral P40 driver 576.57
• CUDA 12.x GPU environment
• Ollama local AI model server
• Open WebUI browser interface
• Local LLM installed and tested
• Both GPUs tested for GPU-accelerated inference
• AI services configured to start with the server
• Model-storage environment configured
• Clean administrator environment supplied
• Temporary credentials provided and should be changed immediately afterreceipt
Tesla Referral P40
Tesla Referral P40
Tesla Referral P40 is a Pascal-generationenterprise datacenter GPU designed for compute workloads.
• Quantization
• Context length
• Software stack
• GPU allocation
• Workload
• Credentials
• User accounts
• Private AI models
• Production databases
• API keys
• Business documents
• Proprietary application / portal code
• Company-specific information
• 2× Intel Xeon Platinum 8160
• 288GB DDR4 ECC RAM
• 2× NVIDIA Tesla Referral P40 24GB
• Dell PERC H730P
• Approximately 17.46TB RAID data volume
• Approximately 1.09TB RAID volume
• 4× 1TB NVMe SSD
• 2× Mellanox ConnectX-4 Lx 25GbE
• 2× Mellanox ConnectX-3 40G
• 2× Dell 1100W redundant PSUs
• iDRAC9 Enterprise
• Windows Server 2025 Datacenter Evaluation installation
• Ollama
• Open WebUI
• Tested public local LLM
• Power cables
• Administrator login information
• iDRAC access information
• Seller data or business information
• Private models / databases
• Proprietary portals or source code
• Rack rails
• Network transceivers
• Network cables
• 288GB ECC memory
• Both Tesla Referral P40 GPUs
• RAID controller
• Installed storage
• NVMe storage
• Network adapters
• Both power supplies
• iDRAC9 Enterprise
