Software Engineer – Full Stack & AI Infrastructure

Thelawofficesofjulietcohenpc — United States · Posted ~22 hours ago

Full-time Onsite $120000-$150000/year

Skills

Full-stack software development AI infrastructure Production systems Open-source AI models Self-hosted inference Secure data systems Full-stack development

🔓 Log in to save this job, tailor your resume & track your apply process — 7 days free, no card needed.

Log in to add to target list

Summary ✨ AI‑Generated

A full-stack software engineering role combining web application development, AI infrastructure, and production engineering. You will translate detailed requirements into reliable software while building infrastructure for specialized AI systems, including open-source model deployment, self-hosted inference, and secure access to sensitive organizational data.

Highlights

Build production-grade full-stack software and AI infrastructure, work closely with technical leadership, and contribute to secure, scalable AI systems with strong ownership across engineering.

Description

Software Engineer Full Stack & AI Infrastructure Firmly · Forest Hills, Queens, NY · Onsite · Full-time $120,000–$150,000/year About Firmly Firmly is building the Law Firm Operating System, an AI-native platform designed to replace the disconnected software, files, and manual workflows law firms rely on today. A major part of the platform is the development of specialized AI systems built for specific legal and operational tasks, with secure access to a firm's documents, data, and internal knowledge. The Role We are looking for an engineer who can work across full-stack development, AI infrastructure, and production systems. You will work directly with the founder, who leads product direction and architecture. Your job is to turn detailed requirements into reliable production software and help build the infrastructure behind Firmly's AI systems. This is not just a web-development role. A major part of the job will involve open-source AI models, self-hosted inference, secure data access, sandboxing, retrieval systems, and AI infrastructure. What You'll Do Build and ship features in our Next.js, React, TypeScript, and PostgreSQL application Deploy and operate open-source LLMs Build specialized AI systems for narrowly defined legal and operational tasks Connect AI securely to firm documents, databases, and file repositories Build RAG, semantic search, embeddings, document ingestion, and retrieval pipelines Design permission-aware AI access to sensitive firm information Build sandboxed environments for AI tools, agents, and generated code Manage containers, model-serving infrastructure, deployments, and production issues Build and maintain authentication, permissions, sessions, APIs, and database systems Review and improve both human-written and AI-generated code Own problems from diagnosis through resolution The Stack Application: Next.js 15 · React 19 · TypeScript · Tailwind · shadcn/ui · PostgreSQL AI / Infrastructure: Open-source LLMs · RAG · Embeddings · Vector Search · Docker · Linux · GPU Inference · Sandboxing · Model Serving The stack is not locked in. We want someone capable of evaluating technologies and making sound architectural decisions. You're a Strong Fit If You Have 3+ years of professional software engineering experience Are strong in TypeScript, React, backend development, and PostgreSQL Are comfortable with Linux, Docker, networking, deployment, and production infrastructure Have experience deploying or working with open-source language models Understand RAG, embeddings, semantic search, and document retrieval Understand authentication, authorization, permissions, and data-access boundaries Can design secure, sandboxed systems for AI execution Can move between frontend, backend, database, infrastructure, and AI work Know how to debug complex systems without saying, “that's not my layer” Use AI coding tools effectively while still understanding and verifying the code they produce Care about shipping software that is secure, tested, reliable, and actually works Especially Valuable Experience with technologies such as vLLM, llama.cpp, Ollama, SGLang, Llama, GLM, Qwen, Mistral, DeepSeek, pgvector, CUDA, GPU infrastructure, model quantization, fine-tuning, or agentic systems is a major plus. Location Full-time, onsite in Forest Hills, Queens, New York. Compensation $120,000–$150,000 per year, depending on experience and demonstrated technical ability.