Vision-Language Models at the Bench: Records Researchers Can Trust
Engineering, IT, Mathematics and Statistics
ABOUT THE INDUSTRY PARTNER
Bower Labs builds Bower, an AI-native lab partner for research scientists. Bower turns lab notes, protocols, meeting recordings, images and video into a structured, searchable scientific record, with an AI assistant that answers questions grounded in a lab’s own data. They also build a smart-glasses client so scientists can capture and consult that record hands-free at the bench.
WHAT’S IN IT FOR YOU?
The intern would partner alongside an engineer on the same problem and their results would directly determine which models are shipped. Bower are open to discussing publication where the work supports it: an anonymised laboratory-protocol video benchmark would be a substantial research artefact in its own right. They do not require prior experience in their specific domain. The intern would be provided with the laboratory context, the data, and an engineer working alongside on the same problem.
RESEARCH TO BE CONDUCTED
The internship would attack this across four connected strands:
- Model landscape and selection. Systematic side-by-side evaluation of current vision-language and vision-language-action models on our own data, weighing accuracy against latency and cost per hour of video.
- Evaluation design. Build a labelled evaluation set from consented laboratory footage and a harness that scores candidate models against it. No public benchmark covers the distribution, so this instrument is what everything else depends on.
- Context engineering. Structured experiments on frame sampling rates, prompt and output schema design, how prior state and protocol documents are supplied to the model, and where quality begins to degrade as context grows.
- Fine-tuning and adaptation. Dataset preparation for, and evaluation of, lightweight adaptation of a smaller model, measured against the baselines established above.
SKILLS WISH LIST
If you’re a postgraduate research student and meet some or all the below we want to hear from you. We strongly encourage women, indigenous and disadvantaged candidates to apply:
Essential:
- Current PhD or Masters by research candidate in computer vision, machine learning or a closely related area.
- Practical experience evaluating models, not only training them. Held-out sets, baselines, ablations, and honest reporting of negative results.
- Strong Python, and comfort working with model APIs and a real code review process.
- Ability to communicate method and findings clearly to an engineering team.
Desirable:
- Work on video understanding: temporal action segmentation, step or procedure recognition, video-language grounding, or egocentric video.
- Familiarity with current vision-language models and how they are evaluated.
- Experience designing an annotation schema and running a labelling effort.
- Exposure to parameter-efficient fine-tuning or adaptation methods.
- Any background in scientific or laboratory workflows.
RESEARCH OUTCOMES
This work evaluated and adapted current VLMs for laboratory procedure understanding across two deployment contexts: real-time bench-side assistance and offline structured record generation.
ADDITIONAL DETAILS
The intern will receive $3,300 per month of the internship, usually in the form of scholarship payments.
It is expected that the intern will primarily undertake this research project during regular business hours and maintain contact with their academic mentor throughout the internship either through face-to-face or phone meetings as appropriate.
The intern and their academic mentor will have the opportunity to negotiate the project’s scope, milestones and timeline during the project planning stage.
Please note, applications are reviewed regularly and this internship may be filled prior to the advertised closing date if a suitable applicant is identified. Early submissions are encouraged.
INTERNSHIP CONTACT
CONNECT WITH APR.INTERN

