Full Stack Software Engineer - AI Infrastructure

Thelawofficesofjulietcohenpc — United States · Posted ~2 hours ago

Mid Full-time Onsite

Skills

full-stack development AI infrastructure production software systems open-source AI models self-hosted inference secure data access JavaScript TypeScript Python AI models APIs

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A growing technology team is seeking a versatile engineer to develop full-stack applications and AI infrastructure. The role involves building reliable production systems, working with modern AI technologies, and transforming complex requirements into scalable software solutions.

Highlights

Opportunity to build advanced AI-powered systems, work across the full software stack, and have significant ownership in designing reliable production infrastructure.

Description

Software Engineer Full Stack & AI Infrastructure Firmly · Forest Hills, Queens, NY · Onsite · Full-time About Firmly Firmly is building the Law Firm Operating System, an AI-native platform designed to replace the disconnected software, files, and manual workflows law firms rely on today. A major part of the platform is the development of specialized AI systems built for specific legal and operational tasks, with secure access to a firm's documents, data, and internal knowledge. The Role We are looking for an engineer who can work across full-stack development, AI infrastructure, and production systems. You will work directly with the founder, who leads product direction and architecture. Your job is to turn detailed requirements into reliable production software and help build the infrastructure behind Firmly's AI systems. This is not just a web-development role. A major part of the job will involve open-source AI models, self-hosted inference, secure data access, sandboxing, retrieval systems, and AI infrastructure. What You'll Do Build and ship features in our Next.js, React, TypeScript, and PostgreSQL application Deploy and operate open-source LLMs Build specialized AI systems for narrowly defined legal and operational tasks Connect AI securely to firm documents, databases, and file repositories Build RAG, semantic search, embeddings, document ingestion, and retrieval pipelines Design permission-aware AI access to sensitive firm information Build sandboxed environments for AI tools, agents, and generated code Manage containers, model-serving infrastructure, deployments, and production issues Build and maintain authentication, permissions, sessions, APIs, and database systems Review and improve both human-written and AI-generated code Own problems from diagnosis through resolution The Stack Application: Next.js 15 · React 19 · TypeScript · Tailwind · shadcn/ui · PostgreSQL AI / Infrastructure: Open-source LLMs · RAG · Embeddings · Vector Search · Docker · Linux · GPU Inference · Sandboxing · Model Serving The stack is not locked in. We want someone capable of evaluating technologies and making sound architectural decisions. You're a Strong Fit If You Have 3+ years of professional software engineering experience Are strong in TypeScript, React, backend development, and PostgreSQL Are comfortable with Linux, Docker, networking, deployment, and production infrastructure Have experience deploying or working with open-source language models Understand RAG, embeddings, semantic search, and document retrieval Understand authentication, authorization, permissions, and data-access boundaries Can design secure, sandboxed systems for AI execution Can move between frontend, backend, database, infrastructure, and AI work Know how to debug complex systems without saying, “that's not my layer” Use AI coding tools effectively while still understanding and verifying the code they produce Care about shipping software that is secure, tested, reliable, and actually works Especially Valuable Experience with technologies such as vLLM, llama.cpp, Ollama, SGLang, Llama, GLM, Qwen, Mistral, DeepSeek, pgvector, CUDA, GPU infrastructure, model quantization, fine-tuning, or agentic systems is a major plus. Location Full-time, onsite in Forest Hills, Queens, New York. Compensation $100,000–$150,000 per year, depending on experience and demonstrated technical ability.