Summary
✨ AI‑Generated
A full-stack software engineering role combining web application development, AI infrastructure, and production engineering. You will translate detailed requirements into reliable software while building infrastructure for specialized AI systems, including open-source model deployment, self-hosted inference, and secure access to sensitive organizational data.
Highlights
Build production-grade full-stack software and AI infrastructure, work closely with technical leadership, and contribute to secure, scalable AI systems with strong ownership across engineering.
Description
Software Engineer Full Stack & AI Infrastructure
Firmly · Forest Hills, Queens, NY · Onsite · Full-time
$120,000–$150,000/year
About Firmly
Firmly is building the Law Firm Operating System, an AI-native platform designed to replace the disconnected software, files, and manual workflows law firms rely on today.
A major part of the platform is the development of specialized AI systems built for specific legal and operational tasks, with secure access to a firm's documents, data, and internal knowledge.
The Role
We are looking for an engineer who can work across full-stack development, AI infrastructure, and production systems.
You will work directly with the founder, who leads product direction and architecture.
Your job is to turn detailed requirements into reliable production software and help build the infrastructure behind Firmly's AI systems.
This is not just a web-development role.
A major part of the job will involve open-source AI models, self-hosted inference, secure data access, sandboxing, retrieval systems, and AI infrastructure.
What You'll Do
Build and ship features in our Next.js, React, TypeScript, and PostgreSQL application
Deploy and operate open-source LLMs
Build specialized AI systems for narrowly defined legal and operational tasks
Connect AI securely to firm documents, databases, and file repositories
Build RAG, semantic search, embeddings, document ingestion, and retrieval pipelines
Design permission-aware AI access to sensitive firm information
Build sandboxed environments for AI tools, agents, and generated code
Manage containers, model-serving infrastructure, deployments, and production issues
Build and maintain authentication, permissions, sessions, APIs, and database systems
Review and improve both human-written and AI-generated code
Own problems from diagnosis through resolution
The Stack
Application:
Next.js 15 · React 19 · TypeScript · Tailwind · shadcn/ui · PostgreSQL
AI / Infrastructure:
Open-source LLMs · RAG · Embeddings · Vector Search · Docker · Linux · GPU Inference · Sandboxing · Model Serving
The stack is not locked in.
We want someone capable of evaluating technologies and making sound architectural decisions.
You're a Strong Fit If You
Have 3+ years of professional software engineering experience
Are strong in TypeScript, React, backend development, and PostgreSQL
Are comfortable with Linux, Docker, networking, deployment, and production infrastructure
Have experience deploying or working with open-source language models
Understand RAG, embeddings, semantic search, and document retrieval
Understand authentication, authorization, permissions, and data-access boundaries
Can design secure, sandboxed systems for AI execution
Can move between frontend, backend, database, infrastructure, and AI work
Know how to debug complex systems without saying, “that's not my layer”
Use AI coding tools effectively while still understanding and verifying the code they produce
Care about shipping software that is secure, tested, reliable, and actually works
Especially Valuable
Experience with technologies such as vLLM, llama.cpp, Ollama, SGLang, Llama, GLM, Qwen, Mistral, DeepSeek, pgvector, CUDA, GPU infrastructure, model quantization, fine-tuning, or agentic systems is a major plus.
Location
Full-time, onsite in Forest Hills, Queens, New York.
Compensation
$120,000–$150,000 per year, depending on experience and demonstrated technical ability.