Senior AI Software Engineer

Nextgen Coding Company — United States · Posted ~3 hours ago

Senior Part-time Hybrid No Visa $20-$25/hour

Skills

Python OCR LLM GPU systems RAG AWS data pipelines GPU

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A senior AI engineering role focused on building intelligent software systems using machine learning, document processing, and large language model technologies.

Highlights

Part-time opportunity to work on advanced AI engineering projects involving large-scale data processing and machine learning technologies.

Description

Education: If you did your undergrad in India, save your time and DO NOT APPLY You will have to submit a W-9 when hired, no OPT, W-8, or other types of sponsorship Senior AI / Software Engineer (Python, OCR, LLM & GPU Systems) NextGen Coding Company Location: New York City — Hybrid / In-Person Required Compensation: $20–$25/hour Engagement Type: Part-Time, Approximately 20 Hours/Week Opportunity: Potential to expand hours based on performance and project needs Work Authorization: Candidates must be legally authorized to work in the United States under an arrangement compatible with a W-9 contractor engagement. NextGen Coding Company cannot provide OPT or employment/visa sponsorship for the role. Role Overview NextGen Coding Company is hiring a highly technical AI / Software Engineer in New York City to help build a large-scale document intelligence and AI platform. The engineer will work heavily with Python, OCR, document processing, data pipelines, LLMs, search/RAG, AWS, and NVIDIA GPU infrastructure. A critical requirement is supporting two versions of the platform: On-prem / air-gapped: running local models such as Qwen on NVIDIA H100/H200 infrastructureCloud: running in AWS with Claude, Qwen, OpenAI, Gemini, and other LLM/model integrationsThe underlying platform must be portable across both environments. Responsibilities Build production Python services, APIs, and data-processing pipelinesProcess large volumes of PDFs, HTML, images, scans, and structured/unstructured dataBuild OCR, extraction, parsing, normalization, and document-classification pipelinesExtract tables, entities, relationships, metadata, citations, and structured recordsBuild ETL and high-volume batch-processing systemsDevelop hybrid search, embeddings, vector search, RAG, reranking, and evidence retrievalIntegrate Claude, Qwen, OpenAI, Gemini, and open-source modelsDeploy and serve local LLMs on NVIDIA H100/H200 GPUsWork with vLLM, PyTorch, Hugging Face, CUDA, quantization, batching, and GPU inference optimizationBuild and maintain the AWS/cloud deploymentWork with PostgreSQL, OpenSearch, S3, Redis, queues, and related data infrastructureBuild APIs, webhooks, Stripe integrations, authentication, and third-party integrationsWrite automated tests covering OCR, extraction, retrieval, data pipelines, and AI outputsDebug complex AI, data, backend, and infrastructure issuesDocument architecture and implementation decisions Required Experience Strong Python engineering experienceStrong backend and data-engineering fundamentalsExperience building production softwareOCR / document intelligence experienceExperience processing PDFs, images, and large datasetsProduction experience with LLMs, RAG, embeddings, and vector searchExperience deploying open-source models on NVIDIA GPU infrastructureUnderstanding of H100/H200-class inference environmentsStrong AWS experiencePostgreSQL / SQLDocker and LinuxREST APIs and third-party integrationsAbility to independently own difficult engineering problems Technical Stack Core: Python, FastAPI, SQL, PostgreSQL, Redis AI / LLM: Qwen, Claude, OpenAI, Gemini, Hugging Face, PyTorch, vLLM OCR / Documents: PaddleOCR, Tesseract, OpenCV, PyMuPDF, Docling or similar Search / Data: OpenSearch/Elasticsearch, pgvector/vector databases, S3/MinIO, ETL pipelines GPU: NVIDIA H100/H200, CUDA, vLLM, model quantization, local inference Cloud: AWS, Bedrock, EC2, EKS/ECS, S3, RDS, OpenSearch, SQS, IAM Infrastructure: Docker, Kubernetes, CI/CD, Linux Product: REST APIs, webhooks, authentication, Stripe, third-party integrations React, Next.js, and TypeScript experience is helpful but not the primary focus. Ideal Candidate We are looking for someone who can be given thousands of pages of PDFs, scans, HTML, images, and other data plus access to an H100/H200 server and an AWS environment — and independently help engineer the system required to process, OCR, structure, search, analyze, and query the information using modern AI models. The ideal candidate is stronger in Python, AI, data processing, OCR, backend systems, and infrastructure than traditional frontend development. You should be comfortable moving between an OCR pipeline, Python data processing, Claude/AWS integration, Qwen running on an H100/H200, databases, APIs, and production infrastructure. About NextGen Coding Company NextGen Coding Company is a U.S.-based software engineering firm building custom software, AI systems, automation platforms, and enterprise applications. Projects span AI/ML, document intelligence, data engineering, financial and compliance technology, cloud infrastructure, and enterprise software. Application Requirements Applicants should provide: Resume or LinkedInGitHub and/or technical portfolioExamples of production AI/data systems builtExperience with Python, OCR, LLMs, AWS, and GPU infrastructureSpecific NVIDIA GPU/model deployment experience, if applicableQualified candidates will participate in a technical discussion and engineering review before engagement.