Software Engineer · AI & Data Systems

Building reliable software for intelligent systems.

I am Anthony Huang, a software engineer with experience in production data infrastructure, distributed systems, machine learning, and LLM applications. I focus on building scalable, observable, and dependable engineering systems.

Professional Profile

Engineering at the intersection of software, data, and AI.

My background combines practical software engineering experience with graduate-level study in artificial intelligence. I have worked on production ETL monitoring, distributed data workflows, model optimisation, LLM-powered analytics, and backend infrastructure.

I build systems with an emphasis on reliability, maintainability, and measurable engineering outcomes. My experience spans cloud data infrastructure, backend services, machine learning experimentation, and deployment-oriented workflows.

I am particularly interested in graduate software engineering opportunities involving backend systems, data platforms, AI infrastructure, distributed computing, and production machine learning.

  • ✓Backend and distributed systems
  • ✓Data infrastructure and production ETL
  • ✓Machine learning and LLM applications
  • ✓Observability, reliability, and deployment

Experience

Applied engineering across production and research environments.

Experience includes production data operations, machine learning research, LLM integrations, cloud data workflows, distributed processing, and backend performance engineering.

  1. Jan 2026 — Present
    Sydney, Australia

    Software Engineering Intern — Data Infrastructure

    BI3 Technologies
    • Monitor production ETL workflows across AWS, Airflow, and SQL Server, validating data freshness, quality, job health, and downstream delivery.
    • Built automated data-quality checks and anomaly-detection workflows with dashboards and alerts, reducing mean time to diagnose pipeline incidents by 60%.
    • Collaborate with engineering and operations teams to investigate production failures and document repeatable remediation steps.
  2. Sep 2024 — Jun 2025
    Hangzhou, China

    Research Assistant

    Zhejiang University — CAD Lab
    • Developed and evaluated PyTorch Diffusion Transformer models for image super-resolution, comparing model variants using high-resolution reconstruction metrics.
    • Profiled latency and memory, then applied structured pruning, quantization, and knowledge distillation, achieving 4× model compression with comparable image quality.
  3. Oct 2023 — Nov 2023
    Beijing, China

    Software Engineering Intern

    Meituan
    • Integrated ChatGLM with an internal analytics platform to translate natural-language questions into SQL workflows for food-operations data.
    • Implemented query validation, real-time retrieval, interactive visualisation, retry logic, and failure recovery for reliable analysis.
  4. Aug 2023 — Sep 2023
    Hangzhou, China

    Software Engineering Intern

    Alibaba Group
    • Built Alibaba Cloud workflows to ingest, transform, and route large-scale IoT telemetry; optimised SQL and batch-processing logic for reliable data flows.
    • Integrated Tmall Genie LLM workflows using LangChain for prompt construction, input preprocessing, model API calls, response handling, and fault recovery.
  5. Aug 2020 — Sep 2020
    Remote

    Software Engineering Virtual Experience Participant

    JPMorgan Chase & Co.
    • Implemented Spark and Kafka transaction-processing pipelines and tuned partitioning and processing logic, reducing benchmark latency by 30%.
  6. Jan 2019 — Feb 2020
    Indonesia

    Software Engineering Intern

    Matrix Mas
    • Developed Linux backend services using Redis caching and asynchronous processing, increasing platform throughput by 10×.

Selected Projects

Technical work demonstrating applied engineering depth.

Projects cover natural language processing, retrieval-augmented generation, distributed systems, and reproducible model deployment.

01

Book Genre Classification & Recommendation System

Python · scikit-learn · NLP · TF-IDF
  • Trained and evaluated Multinomial Naive Bayes, Bernoulli Naive Bayes, Logistic Regression, and Linear SVM text classifiers using cross-validation.
  • Achieved 78.0% accuracy and 0.773 macro-F1 on a five-class classification task.
  • Built a content-based recommender using predicted genres, TF-IDF user profiles, and cosine similarity.
ClassificationRecommendationModel Evaluation
02

RAG Document Intelligence System

Python · Embeddings · Semantic Retrieval · LLMs
  • Built a RAG pipeline for document ingestion, chunking, embedding generation, semantic retrieval, and LLM-based question answering.
  • Implemented top-k retrieval and document summarisation with fallback handling when ranked results were unavailable.
RAGEmbeddingsSemantic Search
03

KeyMesh

C++ · TCP · Multithreading
  • Built a distributed key-value store with a multithreaded TCP server, peer replication, and thread-safe request handling for concurrent clients.
  • Implemented synchronised storage and replication paths to strengthen reliability under concurrent workloads.
Distributed SystemsConcurrencyNetworking
04

LLM Inference & Deployment Platform

Python · Docker · Kubernetes · MLflow
  • Containerised an LLM inference service with health checks, autoscaling, monitoring, experiment tracking, and model versioning.
  • Packaged versioned model artifacts to support repeatable evaluation and rollout across model iterations.
  • Separated model packaging, inference serving, and deployment lifecycle concerns for reproducible releases.
MLOpsKubernetesModel Serving

Technical Skills

A broad foundation across software, data, and machine learning.

Technologies are grouped by practical engineering area to make the profile easy to review for software, infrastructure, and AI-focused roles.

</> Programming Languages

PythonC++JavaSQL

AI Machine Learning

PyTorchscikit-learnClassificationPattern RecognitionNLPComputer VisionData MiningAnomaly DetectionCross-validation

ML Models & Methods

Diffusion TransformersModel IterationPruningQuantizationKnowledge DistillationRAGTF-IDFCosine SimilarityLangChain

SYS Tools & Systems

SparkKafkaAirflowAWSAlibaba CloudSQL ServerPostgreSQLRedisDockerKubernetesMLflowLinux

Education

Computer science foundation with advanced AI study.

Academic training combines computer science fundamentals with graduate study and research in artificial intelligence and synthetic data generation.

University of New South Wales

Master of Information Technology in Artificial Intelligence
Expected May 2027

Research: LLM-driven relational data generation and statistical modelling for synthetic datasets.

Nanjing University

Bachelor of Science in Computer Science
Jun 2024

Foundation in software engineering, algorithms, systems, data, and computer science.

Contact

Open to graduate software engineering opportunities in 2027.

For professional enquiries regarding software engineering, backend systems, data infrastructure, or AI engineering roles, please contact me by email.

anthonyhuang1909@gmail.com